PDFs and photos to CSV
PDF, PNG, JPG and WEBP up to 25 MB. Multi-page PDFs are rendered page by page, every table is extracted and exported as a clean CSV.
Documents in, structured tables out. We started on a Qwen vision model running locally and moved to Gemini 2.5 Flash for production. It reads invoices, bank statements and reports, wrapped in the engineering that makes AI output trustworthy: strict JSON schemas, prompt-injection defenses, validation and billing.
PDF, PNG, JPG and WEBP up to 25 MB. Multi-page PDFs are rendered page by page, every table is extracted and exported as a clean CSV.
Multi-row headers, merged cells, receipts with no grid lines, key–value blocks on invoices. The model first detects the document type, then infers the structure.
Results from each page are stitched back together: headers are normalized so a table that continues on page 3 lands in the same CSV.
A Flutter Web app at app.parseprime.com with drag-and-drop and batch upload, available in 11 interface languages. A separate Flutter Web admin panel for operations.
Business accounts get API keys and a documented REST API, so extraction can run inside their own data pipelines.
Accounts, credits and Stripe checkout, GDPR-compliant terms, operated by TAGonSoft SRL at parseprime.com.
Calling a model is the easy part. These are the pieces that turn it into a product people pay for.
Low temperature and JSON-only responses against a fixed schema (headers, rows, meta with document type and confidence). Anything that does not parse is rejected, never shown.
Uploaded documents are untrusted input. Text inside a file is always treated as data: instruction-like content is extracted literally and flagged, never obeyed.
Each plan maps to a fixed reasoning (thinking) and output token budget, so higher tiers get more careful extraction while the cost per page stays predictable.
CSV output is sanitized against spreadsheet formula injection, downloads use signed, expiring links, and transient AI errors are retried automatically.
The AI service sits behind an interface with a stub for the test suite, and a health endpoint checks the AI provider alongside the database.
A full product stack, designed and delivered by our team with AI-assisted development.
🧠 AI
Qwen (local, Ollama) → Gemini 2.5 Flash
Prototyped on a self-hosted Qwen vision model via Ollama, then moved to Gemini 2.5 Flash for accuracy and scale. The AI layer sits behind one interface, so the model can change without touching the product. Prompting, guardrails and post-processing are ours.
🔧 Backend
PHP 8.4 + Symfony 7
REST API with JWT and API-key auth, credits, Stripe webhooks and a background worker for batch jobs. PDFs are rendered with Poppler.
🖥️ Web apps
Flutter Web
The customer app and the admin panel, both in Flutter Web, plus a static marketing site and API docs.
💾 Data & infra
PostgreSQL · Docker · Hetzner
Docker Compose behind Nginx with TLS, scripted deploys and one-command rollback.
We integrate LLMs where they create real value — with the validation, security and cost control that production needs. Tell us about your documents, data or workflow.