- 01Clean PDF
- 02OCR + parser
- 03Evidence lock
- 04Human review
Source P&ID drawings are intentionally not published.
Overview
Objective and scope
The workflow reads a clean P&ID during inference, then preserves raw alternatives, evidence crops, prediction locks, and post-lock scoring. It is intentionally bounded to line-label localization and transcription; it does not infer piping connectivity or deliver a final production line register.
Engineering question
Manual line-label extraction is repetitive, but a seemingly plausible OCR reading can be wrong or out of drawing scope. The pipeline therefore keeps detection, parsing, review, and evaluation traceable rather than hiding uncertainty behind automatic acceptance.
Objectives
- Locate and transcribe candidate line identifiers from a clean P&ID without pre-reading held-out truth.
- Retain evidence for candidate selection, alternatives, crops, confidence flags, and post-lock evaluation.
- Keep a human review gate in front of any register use.
Method
How the work was approached
- Run dual-orientation OCR with high-resolution local refinement and restore candidate geometry to drawing coordinates.
- Apply loose detection, strict parsing, candidate fusion, and review-status flags without accepting records automatically.
- Lock held-out predictions before marked-PDF, Excel, or adjudicated-truth access, then score localization and transcription separately.
Tools
Assumptions and boundaries
- The benchmark is for line-label localization and transcription only.
- Confirmed visible-label extents are evaluated separately from eight ambiguous occurrences that need more adjudication.
- AUTO is disabled; auto-eligible is diagnostic only and accepts no production record.
Workflow illustration derived from the retained benchmark protocol
From evidence to review
- 01Clean PDF
- 02Dual-orientation OCR
- 03Parser and fusion
- 04Evidence crops
- 05Prediction lock
- 06Human review
- 07Post-lock evaluation
Evidence
What the public case is based on
- V1 README, artifact inventory, and pipeline modules for OCR, parsing, fusion, geometry, inference, and validation.
- Held-out benchmark protocol and validation report.
- Evaluator V1.2 audit and comparison, including its further-adjudication decision.
Validation
- The validation report confirms held-out prediction locks preceded first ground-truth access and all 25 unit/regression tests passed.
- The evaluator performs spatial pairing before reading strings, preventing transcription text from driving localization matches.
- The current strict benchmark still requires further adjudication because eight visible-label boxes remain ambiguous.
Results
Selected, bounded observations
- 57 / 65
- confirmed-label localization87.69%; strict current evaluator
- 89.47%
- drawing-wide precisionscope-adjudication based
- 0
- auto-accepted recordsAUTO intentionally disabled
| Control | Observed evidence | Why it matters |
|---|---|---|
| Blind protocol | Held-out predictions locked before truth access | Prevents post-hoc tuning against the evaluated drawing. |
| Localization | 57/65 confirmed labels localized; eight remain ambiguous | The benchmark needs further adjudication before high-stakes use. |
| Acceptance gate | AUTO disabled; 0 records accepted automatically | Human review remains required for every usable output. |
Reported results
- The latest strict confirmed-label evaluation reports 57/65 localization recall (87.69%).
- Drawing-wide precision is 89.47%; the score includes visible valid line labels, duplicates, and scope classification rather than treating every prediction as a register row.
- AUTO acceptance remains disabled with 0 accepted records, even where diagnostic auto-eligibility was observed.
Key takeaways
- Confidence must not substitute for review when OCR alternatives, nearby labels, or broad clusters can change the match.
- A useful engineering-automation MVP can reduce review effort while still making every acceptance decision inspectable.
Blind benchmark for line-label localization and transcription. The specific limitations below remain part of the case, not footnotes.
Limitations
- This is an assisted-review benchmark, not an autonomous line-register system. FROM/TO, connectivity, piping graph inference, production numbering, and final Excel output are outside scope.
- The current strict benchmark requires further adjudication of eight ambiguous visible-label occurrences before a high-stakes performance claim would be appropriate.
What I learned
- Lock predictions before opening truth so evaluation remains auditable.
- Keep geometric localization and text scoring separate, then preserve raw evidence for a reviewer.