Document extraction with eval-gated CI

Problem → approach → proof, then open the live sibling demo

DocExtract AI is a sibling portfolio module (docextract). It is not built inside this EnterpriseHub monorepo and is not served by an EnterpriseHub API. This chapter is narrative + static media + outbound CTAs only.

  • Problem: messy PDFs and forms need structured fields with citations, not chatty ungrounded summaries.
  • Approach: FastAPI extraction pipeline, cost-aware model routing, independent LLM judge, and an eval gate that fails PRs on accuracy regression.
  • Proof split: docextract owns doc-RAG evidence; EnterpriseHub /quality owns platform evals, OTel, and A/B artifacts.

Streamlit Cloud may cold-start; if the live demo sleeps, use the GitHub fallback. EH CI does not assert external Streamlit uptime.

In-app preview

Static screenshot from the sibling repo — no iframe

DocExtract AI live demo: document extraction with evaluation scores, agent trace, and cost analysis

Provenance: frontend/public/intelligence/PROVENANCE.md. Copied from docextract docs/screenshots/demo-hero.png on 2026-07-12.