Document extraction with eval-gated CI
Problem → approach → proof, then open the live sibling demo
DocExtract AI is a sibling portfolio module (docextract). It is not built inside this EnterpriseHub monorepo and is not served by an EnterpriseHub API. This chapter is narrative + static media + outbound CTAs only.
- Problem: messy PDFs and forms need structured fields with citations, not chatty ungrounded summaries.
- Approach: FastAPI extraction pipeline, cost-aware model routing, independent LLM judge, and an eval gate that fails PRs on accuracy regression.
- Proof split: docextract owns doc-RAG evidence; EnterpriseHub /quality owns platform evals, OTel, and A/B artifacts.
Streamlit Cloud may cold-start; if the live demo sleeps, use the GitHub fallback. EH CI does not assert external Streamlit uptime.
In-app preview
Static screenshot from the sibling repo — no iframe

Provenance: frontend/public/intelligence/PROVENANCE.md. Copied from docextract docs/screenshots/demo-hero.png on 2026-07-12.