Returns analytics: where the product does not sell
A supplier ships goods to national retail chains and receives write-offs back — each chain in its own format: Excel exports, PDF return notices, scans and photos of documents. Records were kept by hand in a dozen spreadsheets, and nobody calculated the return rate per store. The task was a system that collects data from heterogeneous sources by itself and shows where the product does not sell.
What is inside
- An adapter for every counterparty. Report formats differ fundamentally — daily detail at one chain, a weekly pivot at another. The file type is detected by content, not by name.
- A PDF recognition pipeline. Digital documents are read through the text layer, scans go to OCR with automatic rotation detection, and disputed cases are finished by the model's vision. Multi-page files are parsed page by page, picking the right document among invoices, settlement forms and signature sheets.
- Protection from false data. The sum of line items is reconciled with the document total. If it does not match, the record is blocked and the document goes to manual review instead of silently entering made-up figures.
- Reversible operations. Every load keeps a snapshot of the state before and after; any load can be excluded from calculations and brought back without losing the rest of the data.
- Analytics. Return rate per store and product pair, drill-down from chain to store to item, totals in money and in units, completeness control of loaded data, a report on high-loss items.
Quality control
Development ran in an edit — test — live check — deploy loop. Every change was closed with an automated test, and visual changes were checked in the browser.
- 376 automated tests in pytest: business logic, parsers, API.
- 12 browser tests that execute the real frontend code in an isolated environment with a substituted DOM.
- Three-level isolation of the test database, ruling out writes to production data.
- An integrity check after every deploy: file checksums, responses of all endpoints, metrics compared before and after.
A separate class of defects was caught only by live checks: cascading CSS conflicts, frozen animations, elements present in the DOM but invisible on screen.
Stack
- Backend: Python 3, FastAPI, SQLAlchemy, SQLite ready for PostgreSQL, Uvicorn.
- Data processing: pandas, openpyxl, an ETL pipeline with idempotent loading.
- Document recognition: pdfplumber, PyMuPDF, Tesseract OCR, Claude Vision API.
- Frontend: vanilla JavaScript, Jinja2, CSS with no frameworks and no build step.
- Testing: pytest, Node.js and vm for DOM tests.
- Infrastructure: systemd, rsync deploy, basic-auth, snapshots with rollback.
Outcome
- Manual data collection replaced by automatic loading from every source.
- The return rate is calculated for every store and every item.
- The share of documents needing manual entry cut fourfold through improvements to the recognition pipeline.
FAQ
How does the system know what kind of file it is?
What happens if recognition gets it wrong?
Can a failed load be undone?
Why browser tests when there are 376 pytest tests?
Tell us about your task
We reply within one business day and send an estimate in 1–2 days. Free of charge.