Maritime document intelligence · Evaluation note
How to evaluate maritime document intelligence
Published 11 September 2026
A document system is not trustworthy because a demonstration looks fluent. Evaluation must test field-level extraction, source traceability, rule coverage, unresolved exceptions and the human workload needed to reach an accountable decision.
Begin with the decision boundary
The pilot maps agreed maritime documents into a customer-defined schema. It does not autonomously hire, reject or approve seafarers. Consequential decisions, missing documents, conflicts and low-confidence fields remain with authorized people.
A measurable acceptance set
Representative sample
Agree document types, languages, quality ranges and exception cases before evaluation.
Ground truth
Create a human-reviewed reference set with explicit field definitions and source locations.
Field accuracy
Measure correct values, omissions and unsupported additions at the field level.
Traceability
Check that each consequential field points back to the exact supporting evidence.
Rules and exceptions
Measure rule coverage, conflicts, confidence thresholds and the review workload they create.
Operational acceptance
Test security, access, auditability, integrations and human approval inside the customer boundary.
What the pilot must report
Results must name the sample, metrics, environment, limitations and unresolved exceptions. Aggregate accuracy alone is insufficient when a low-frequency field can change a consequential decision. Changes to schemas, rules or document mix require re-evaluation.
Standards and governance context
The evaluation approach is consistent with the NIST AI Risk Management Framework’s focus on trustworthiness across design, development, use and evaluation. Maritime rules and certification decisions remain governed by applicable authorities and the customer’s approved procedures, including the STCW framework where relevant.