AI Document Intelligence Platform
Production-grade AI system for automatic document classification, extraction, and validation
Overview
Problem
Enterprises process thousands of documents daily - invoices, contracts, compliance forms, receipts - each with varying formats, structures, and quality. Manual data entry is slow and error-prone. Traditional OCR and template-based extraction fail when documents don't match expected layouts. Different vendors use different field names, date formats, currencies, and layouts for the same information type.
Solution
Built an AI-powered document processing pipeline that combines semantic understanding with deterministic validation. The system classifies documents hierarchically, selects appropriate extraction schemas, uses LLMs for semantic field extraction, normalizes data across format variations, validates with business rules, calculates confidence scores, and routes uncertain cases to human review. Designed to achieve ≥98% extraction accuracy on supported document types.
Architecture

The platform follows a multi-stage pipeline architecture: Upload → Text Extraction/OCR → Document Classification → Schema Selection → LLM Structured Extraction → Pydantic Validation → Business Validation → Confidence Scoring → Human Review (if needed) → PostgreSQL. Each stage is modular, traceable, and recoverable from failures.
Document Ingestion - Handles uploads, stores originals, detects text vs scanned documents
Text Extraction Pipeline - PDF text extraction and OCR for scanned documents
Classification Engine - Hierarchical classifier (family → specific type) with confidence thresholds
Schema Manager - Pydantic schemas for each document type with validation rules
LLM Extraction Service - OpenAI structured outputs for semantic field extraction
Validation Layer - Deterministic business rules (e.g., subtotal + tax ≈ total)
Confidence Scoring - Field-level and document-level confidence calculation
Human Review Queue - Interface for reviewing low-confidence extractions
Evaluation Framework - Measures accuracy against labeled test datasets
Observability Dashboard - Processing metrics, accuracy, costs, and error tracking
Key Engineering Decisions
Challenge
Documents have same semantic meaning but different formats (e.g., 'Total' vs 'Amount Due' vs 'Grand Total')
Decision
Use semantic extraction instead of coordinate-based template matching
Reasoning
LLMs understand semantic equivalence and can extract 'total' regardless of label variations. This works for non-standardized documents where field positions vary. Template matching would require maintaining separate extractors for every vendor format.
Alternatives Considered
- •Template matching per vendor - brittle, high maintenance, fails on new formats
- •Keyword search - misses semantic variations and context-dependent fields
Challenge
LLM extraction can hallucinate or produce invalid data structures
Decision
Implemented multi-layer validation: Pydantic type validation → deterministic business rules → confidence thresholds
Reasoning
Pydantic enforces schema structure at runtime. Business rules catch semantic errors (e.g., due_date before issue_date). Confidence thresholds route uncertain cases to human review. Never trust LLM output blindly - always validate with deterministic logic when possible.
Alternatives Considered
- •Trust LLM output directly - dangerous, hallucinations reach production
- •Rules-only validation - misses complex semantic relationships
Challenge
Need to measure actual accuracy rather than claiming arbitrary performance
Decision
Built evaluation framework with labeled test dataset measuring field-level precision, recall, and F1 scores
Reasoning
Evaluation against ground truth reveals which document types meet the ≥98% accuracy target and which need improvement. Tracks exact match rate, confidence calibration, and human review rate. Enables data-driven optimization instead of guessing.
Challenges & Solutions
Problem
Non-standardized documents have the same data in completely different structures (tables vs paragraphs, different field names, mixed formats)
Solution
Designed for semantic extraction from the start. LLM receives full document context and understands that 'Client', 'Customer', 'Bill To' can represent the same field. Implemented normalization layer that converts various date formats (DD/MM/YYYY vs MM-DD-YY), currency representations ($1,234.50 vs 1.234,50), and language variations into unified schema.
Outcome
System handles documents from multiple vendors without per-vendor configuration. Normalization ensures downstream systems receive consistent data regardless of source format.
Problem
Some fields have high confidence but are semantically incorrect (e.g., extracting 'issue date' as 'due date')
Solution
Implemented deterministic validation rules that check business logic constraints (due_date ≥ issue_date, subtotal + tax ≈ total, tax_id format validation). When validation fails, document is flagged for review even if extraction confidence was high. Store source text references for every field to enable human verification.
Outcome
Catches semantic errors that confidence scores miss. Traceability allows reviewers to quickly verify extraction against source document. Reduces false positives in automatic approval.
What I Learned
- →Prefer deterministic logic over LLMs whenever possible - use AI for semantic understanding, not arithmetic
- →Every extracted field needs source traceability - especially critical for legal documents
- →Confidence scores must be calibrated - a 95% confidence that's actually 70% accurate is worse than honest uncertainty
- →Human-in-the-loop is not a fallback, it's a core feature - design the review interface from day one
- →Real accuracy measurement is hard but essential - labeled test datasets reveal what actually works
- →Document quality varies wildly - plan for OCR errors, missing fields, and non-standard layouts
- →Normalization is as important as extraction - inconsistent data formats break downstream systems