Most organisations do not have a document problem. They have an operational problem caused by information trapped in documents, emails, reports and free text.
Start with the target schema
Before choosing a model, define what the business needs: entities, fields, relationships, confidence requirements and provenance. A strong schema turns an AI problem into a measurable data-quality problem.
Design a pipeline, not a prompt
Production needs ingestion, parsing, classification, extraction, validation, enrichment, exception handling and storage. Some steps use language models; others should be deterministic. Every step should be observable.
Put humans where errors cost most
Route low-confidence or high-consequence cases to reviewers and let predictable cases flow. Capture the corrections: they become evaluation data and, often, training data.
Measure at the field and workflow level
One global accuracy number is rarely enough. Track performance by document type, field and consequence, and measure cycle time, review burden and downstream failure rates.
The best document-AI systems are not the ones with the cleverest prompt. They have a clear schema, an observable pipeline and a deliberate error strategy.