A Binary Answer Is Not a Simple Decision: Testing a Specialist AI Model
An eight-case comparison of a specialist decision model and general models: expert agreement, missing answers, and what the test does not establish.
The answer may be only yes or no. Choosing it can still require expert judgment.
A model specialised for decisions is an interesting candidate for that work. But does its specialisation help when the options are simple and the circumstances are ambiguous?
Agent Reliability Lab explored that question on 4 October 2026: a local decision model, Winnow-12B, compared with two general models on eight cases judged by an experienced decision owner. The test produced useful observations, but it did not isolate the effect of specialisation. There was no comparison with Winnow’s base model, and its dedicated decision interface was not tested.
Two options, no ready-made formula
Two options. No ready-made formula.
Remaining defects
Cost of delay
Limits of rollback
Interpret the circumstances and conflicting priorities.
No fixed formula settles every case.That is the type of judgment the study examined. It was not a test of whether a model could execute a fully specified rule. The actual specialist domain, detailed inputs and decision factors remain private; the release example illustrates the distinction without replacing the original cases.
What was compared
There were eight related, semi-synthetic cases. Historical observations were combined with deliberately varied conditions, and each case offered two possible decisions.
The expert made the reference decisions without seeing the model answers, as the participant directly confirmed on 7 October. The expert reported recognising none of the underlying historical cases.
The reference measured agreement with that expert, not independently established correctness or business outcomes. A model can disagree with a person without this test proving that either answer is objectively wrong.
Each case was presented in three ways, as separate requests without conversation history:
- Decide directly: choose without hints.
- Consider both sides: give arguments for and against, then decide.
- Use supplied expert factors: weigh two factors stated by the expert, then decide.
The third approach changes the question. It tests how the model applies supplied factors, not whether it finds those factors unaided.
The comparison included Claude Sonnet 5.5 and Claude Opus 5.5, both at medium effort, and Winnow-12B through ordinary chat with reasoning enabled and disabled. The hosted-model versions and effort setting were confirmed by the experiment author; complete API request metadata was not retained.
The recorded results
Agreement with the expert
| Configuration | Decide directly | Consider both sides | Use expert factors |
|---|---|---|---|
| Claude Sonnet 5.5, medium | 5/8 | 6/8 | 7/8 |
| Claude Opus 5.5, medium | 5/8 | 3/8 | 4/8 |
| Winnow-12B, reasoning enabled | 3/8 | 3/8 | 2/8 |
| Winnow-12B, reasoning disabled | 6/8 | 7/8 | 4/8 |
Each column repeats the same eight cases: these are not 24 independent cases. The expert chose one option six times and the other twice. No ranking or significance claim.
That baseline matters for interpreting Winnow’s direct-decision result with reasoning disabled. It returned the same choice and confidence on all eight cases. Its 6/8 score did not discriminate between them. This observation does not establish which inputs the model processed internally.
Two configurations reached 7/8: Sonnet using expert factors and Winnow without reasoning when asked to consider both sides. Both exceeded the baseline by one case. Neither result establishes superiority on future decisions, and the scores do not demonstrate that either model weighed the important facts correctly.
Supplying expert factors did not improve every configuration. Winnow without reasoning scored 7/8 when considering both sides and 4/8 when given the factors. On these cases, the prompt approach mattered; there is no general instruction to add expert factors and expect an improvement.
Sometimes the model did not deliver a decision
Reasoning enabled: 24 requests
Delivered: 10/24. Agreement among delivered decisions: 8/10.
8/10 is not overall reliability across all 24 requests.
This shows a delivery problem in the tested configuration. It does not show that reasoning is generally harmful or that a dedicated decision interface would behave the same way.
What the specialist-model test tells us
The local specialist was not simply unable to match the expert: one of its prompt configurations reached 7/8, the same count as the best recorded general-model configuration. It also produced a constant-answer result and substantial missing output in other configurations.
Those observations are more useful than a blanket verdict that a specialist model works or does not work. They identify where this setup deserves a closer look.
They do not measure the benefit of specialisation itself. Without the corresponding base-model control, the comparison cannot attribute a result to decision-focused training. And Winnow’s model card describes a dedicated interface that scores predefined options; this study used chat instead.
Eight related cases are also too few to establish a model ranking. Several prompts were compared on the same examples, and no fresh hold-out cases were tested. The expert’s agreement is the reference for this study, not a universal definition of the correct decision.
Choose the next test by the question
Different questions need different tests
Dedicated decision endpoint
Chat was tested; the dedicated interface was not.Matched base-model control
No base-model comparison was run.Fresh blinded cases
No fresh hold-out set was tested.For an operator, the practical conclusion is to separate three things: the simplicity of the answer format, the judgment required by the task, and the evidence supporting a particular model configuration. A binary answer does not make the second or third simple.
Method and limits
The study used eight cases and three requests per case for each configuration on 4 October 2026. Saved answers and expert records reproduce all twelve agreement counts. The original committed snapshot confirms 14 empty final answers; an earlier narrative count of 13 was incorrect.
The local model was Winnow-12B in Q8 form, served through llama.cpp chat. Its saved runner specifies temperature zero and a 6,000-token output budget. The handoff records build b10580 and the reasoning-off flag, but complete request metadata does not independently certify those runtime settings. Structured API finish reasons were not saved.
The author identified Sonnet 5.5 and Opus 5.5, both medium, on 7 October. Immutable hosted-model IDs were not retained. The expert directly confirmed not seeing the model answers when making the reference decisions; that does not assert that reference files were saved before model generation.
To examine a bounded decision workflow, start with a written brief: the two options, the evidence that matters, and the authority that should remain with a person.
- Separate a short answer format from the judgment needed to choose it.
- State whether the reference measures expert agreement or proven outcomes.
- Compare agreement with a constant-answer baseline.
- Keep missing final answers separate from disagreeing decisions.
- Match any follow-up test to the specific uncertainty it should resolve.