Skip to content
Welf LabsResearch for industrial intelligence.

Test protocol · no measured results

One industrial workflow, two models: A protocol for testing a model switch

Founder & CEO of Welf3 min read
Process overview: One industrial workflow, two models: A protocol for testing a model switch

A model replacement succeeds when the workflow still produces an acceptable result. This protocol proposes a limited comparison: two models extract technical requirements from the same synthetic requests and produce a draft for review.

Status: test protocol. This experiment has no measured results yet. The case counts and repetitions below are proposed design choices. They are not an industry standard or evidence from a Welf customer deployment.

Download the blank results template (CSV)

Define the question before running the test

Can the application use a second model without changing its business rules or approvals, while still producing a useful requirements draft?

The application must return structured fields, source references, contradictions and missing information. It must not set prices or write to a production ERP. The comparison covers output quality, reviewer effort and the changes required to switch.

Prepare a bounded evaluation set

The proposed set contains 40 synthetic requests for fictional parts: 20 ordinary complete cases, ten with missing or contradictory information, five with format or revision issues, and five containing unauthorised instructions in source documents.

Keep development examples separate. A domain specialist defines the expected result and a second reviewer checks disputed interpretations. Synthetic cases can still be unrealistic or too easy; their construction needs scrutiny.

Record the model identifiers, available version information, parameters, prompt version, environment and date. Both variants use the same evaluation material and business rules.

Separate direct replacement from adaptation

Run one comparison with the application held as constant as possible. Run a second only if documented changes are needed to make the alternatives useful. Keep those findings distinct. A tuned deployment and a direct replacement answer different questions.

For each variant, record:

MeasureDefinition
Fully correct draftRequired fields, evidence and unresolved points are all correctly represented
Critical errorInvented required value, missed defined contradiction or unauthorised action
Reviewer effortMeasured time to check and correct the result
RuntimeEnd-to-end duration, including slower cases
CostCalls, infrastructure and review with stated accounting assumptions
Switching effortTime spent and components changed

The proposal uses three runs per model and case: 240 executions across two models. There are still only 40 distinct cases. Repetition does not create broader task coverage.

Keep the assessment inspectable

Where practical, conceal model identity from reviewers and vary result order. Record disagreements and show examples of failures. A separate integration test should examine timeouts and invalid formats; do not present its outcome as a measure of model intelligence.

Existing approvals remain in place even if every evaluation case succeeds. The experiment cannot establish production reliability or complete a system safety assessment.

A later results release must identify the actual configuration, dataset rights, scoring rules, absolute counts and limitations. Until the work is performed, this protocol supports no winner or savings claim.

Welf Labs connects this evaluation work to industrial applications. For a production decision, adapt the protocol to your own workflow and its consequences.

Plan a model comparison with Welf