2024•Case Study

AI Document Intelligence Platform

Production-grade AI system for automatic document classification, extraction, and validation

PythonFastAPIPydanticOpenAI APIPostgreSQLReactTypeScriptTailwind CSS

Overview

Problem

Enterprises process thousands of documents daily - invoices, contracts, compliance forms, receipts - each with varying formats, structures, and quality. Manual data entry is slow and error-prone. Traditional OCR and template-based extraction fail when documents don't match expected layouts. Different vendors use different field names, date formats, currencies, and layouts for the same information type.

Solution

Built an AI-powered document processing pipeline that combines semantic understanding with deterministic validation. The system classifies documents hierarchically, selects appropriate extraction schemas, uses LLMs for semantic field extraction, normalizes data across format variations, validates with business rules, calculates confidence scores, and routes uncertain cases to human review. Designed to achieve ≥98% extraction accuracy on supported document types.

Architecture

AI Document Intelligence Platform Architecture

The platform follows a multi-stage pipeline architecture: Upload → Text Extraction/OCR → Document Classification → Schema Selection → LLM Structured Extraction → Pydantic Validation → Business Validation → Confidence Scoring → Human Review (if needed) → PostgreSQL. Each stage is modular, traceable, and recoverable from failures.

Document Ingestion - Handles uploads, stores originals, detects text vs scanned documents

Text Extraction Pipeline - PDF text extraction and OCR for scanned documents

Classification Engine - Hierarchical classifier (family → specific type) with confidence thresholds

Schema Manager - Pydantic schemas for each document type with validation rules

LLM Extraction Service - OpenAI structured outputs for semantic field extraction

Validation Layer - Deterministic business rules (e.g., subtotal + tax ≈ total)

Confidence Scoring - Field-level and document-level confidence calculation

Human Review Queue - Interface for reviewing low-confidence extractions

Evaluation Framework - Measures accuracy against labeled test datasets

Observability Dashboard - Processing metrics, accuracy, costs, and error tracking

Key Engineering Decisions

Challenge

Documents have same semantic meaning but different formats (e.g., 'Total' vs 'Amount Due' vs 'Grand Total')

Decision

Use semantic extraction instead of coordinate-based template matching

Reasoning

LLMs understand semantic equivalence and can extract 'total' regardless of label variations. This works for non-standardized documents where field positions vary. Template matching would require maintaining separate extractors for every vendor format.

Alternatives Considered

  • •Template matching per vendor - brittle, high maintenance, fails on new formats
  • •Keyword search - misses semantic variations and context-dependent fields

Challenge

LLM extraction can hallucinate or produce invalid data structures

Decision

Implemented multi-layer validation: Pydantic type validation → deterministic business rules → confidence thresholds

Reasoning

Pydantic enforces schema structure at runtime. Business rules catch semantic errors (e.g., due_date before issue_date). Confidence thresholds route uncertain cases to human review. Never trust LLM output blindly - always validate with deterministic logic when possible.

Alternatives Considered

  • •Trust LLM output directly - dangerous, hallucinations reach production
  • •Rules-only validation - misses complex semantic relationships

Challenge

Need to measure actual accuracy rather than claiming arbitrary performance

Decision

Built evaluation framework with labeled test dataset measuring field-level precision, recall, and F1 scores

Reasoning

Evaluation against ground truth reveals which document types meet the ≥98% accuracy target and which need improvement. Tracks exact match rate, confidence calibration, and human review rate. Enables data-driven optimization instead of guessing.

Challenges & Solutions

Problem

Non-standardized documents have the same data in completely different structures (tables vs paragraphs, different field names, mixed formats)

Solution

Designed for semantic extraction from the start. LLM receives full document context and understands that 'Client', 'Customer', 'Bill To' can represent the same field. Implemented normalization layer that converts various date formats (DD/MM/YYYY vs MM-DD-YY), currency representations ($1,234.50 vs 1.234,50), and language variations into unified schema.

Outcome

System handles documents from multiple vendors without per-vendor configuration. Normalization ensures downstream systems receive consistent data regardless of source format.

Problem

Some fields have high confidence but are semantically incorrect (e.g., extracting 'issue date' as 'due date')

Solution

Implemented deterministic validation rules that check business logic constraints (due_date ≥ issue_date, subtotal + tax ≈ total, tax_id format validation). When validation fails, document is flagged for review even if extraction confidence was high. Store source text references for every field to enable human verification.

Outcome

Catches semantic errors that confidence scores miss. Traceability allows reviewers to quickly verify extraction against source document. Reduces false positives in automatic approval.

What I Learned

  • →Prefer deterministic logic over LLMs whenever possible - use AI for semantic understanding, not arithmetic
  • →Every extracted field needs source traceability - especially critical for legal documents
  • →Confidence scores must be calibrated - a 95% confidence that's actually 70% accurate is worse than honest uncertainty
  • →Human-in-the-loop is not a fallback, it's a core feature - design the review interface from day one
  • →Real accuracy measurement is hard but essential - labeled test datasets reveal what actually works
  • →Document quality varies wildly - plan for OCR errors, missing fields, and non-standard layouts
  • →Normalization is as important as extraction - inconsistent data formats break downstream systems