Test protocol · no measured results
One industrial workflow, two models: A protocol for testing a model switch
A model replacement succeeds when the workflow still produces an acceptable result. This protocol proposes a limited comparison: two models extract technical requirements from the same synthetic requests and produce a draft for review.
Status: test protocol. This experiment has no measured results yet. The case counts and repetitions below are proposed design choices. They are not an industry standard or evidence from a Welf customer deployment.
Download the blank results template (CSV)
Define the question before running the test
Can the application use a second model without changing its business rules or approvals, while still producing a useful requirements draft?
The application must return structured fields, source references, contradictions and missing information. It must not set prices or write to a production ERP. The comparison covers output quality, reviewer effort and the changes required to switch.
Prepare a bounded evaluation set
The proposed set contains 40 synthetic requests for fictional parts: 20 ordinary complete cases, ten with missing or contradictory information, five with format or revision issues, and five containing unauthorised instructions in source documents.
Keep development examples separate. A domain specialist defines the expected result and a second reviewer checks disputed interpretations. Synthetic cases can still be unrealistic or too easy; their construction needs scrutiny.
Record the model identifiers, available version information, parameters, prompt version, environment and date. Both variants use the same evaluation material and business rules.
Separate direct replacement from adaptation
Run one comparison with the application held as constant as possible. Run a second only if documented changes are needed to make the alternatives useful. Keep those findings distinct. A tuned deployment and a direct replacement answer different questions.
For each variant, record:
| Measure | Definition |
|---|---|
| Fully correct draft | Required fields, evidence and unresolved points are all correctly represented |
| Critical error | Invented required value, missed defined contradiction or unauthorised action |
| Reviewer effort | Measured time to check and correct the result |
| Runtime | End-to-end duration, including slower cases |
| Cost | Calls, infrastructure and review with stated accounting assumptions |
| Switching effort | Time spent and components changed |
The proposal uses three runs per model and case: 240 executions across two models. There are still only 40 distinct cases. Repetition does not create broader task coverage.
Keep the assessment inspectable
Where practical, conceal model identity from reviewers and vary result order. Record disagreements and show examples of failures. A separate integration test should examine timeouts and invalid formats; do not present its outcome as a measure of model intelligence.
Existing approvals remain in place even if every evaluation case succeeds. The experiment cannot establish production reliability or complete a system safety assessment.
A later results release must identify the actual configuration, dataset rights, scoring rules, absolute counts and limitations. Until the work is performed, this protocol supports no winner or savings claim.
Welf Labs connects this evaluation work to industrial applications. For a production decision, adapt the protocol to your own workflow and its consequences.