Broker submissions arrived as 60-page PDF bundles. We built a retrieval and extraction pipeline that turns each bundle into a structured risk file, with every field traceable to its source page.
Underwriters spent 40% of their day re-keying data from broker PDFs into a legacy rating engine. Accuracy mattered more than speed — a wrong sum insured is a regulatory problem, not a UX one.
Built a document pipeline splitting bundles into typed sections before extraction
Grounded every extracted field with a page-and-bounding-box citation
Added a confidence gate that routes uncertain fields to human review
Shipped a review UI so underwriters correct rather than re-key
73%
less manual data entry
4.5×
faster submission turnaround
99.2%
field-level extraction accuracy
"They refused to ship the model until the citation trail worked. That call is the reason our compliance team signed off in one review."