HowardBench · completed 1 September 2026
We asked the same model the same mechanical-engineering questions twice: once alone, and once with Howard’s engineering calculations and reviewed evidence available through MCP.
Raw, uncorrected scorer results. Howard did not pass the holdout release gates: it missed required routes and produced 33 numeric entries absent from its supporting trace. Both are published below.
Both arms used the same model identifier and the same benchmark system prompt. The unaided arm had no tools, web access, or files. The Howard arm could call only the authenticated 45-capability Howard MCP.
| Suite | Without Howard | With Howard | Improvement |
|---|---|---|---|
| 675 tool questions | 23.2% | 66.9% | +43.6 points |
| 30 design problems | 8.3% | 63.3% | +55.0 points |
| Metric | Without Howard | With Howard |
|---|---|---|
| Task accuracy | 23.2% | 66.9% |
| Full-item pass | 2.5% | 57.6% |
| Grounded-number accuracy | 0.0% | 51.2% |
| Correct refusal | 44.7% | 73.9% |
| Choice accuracy | 16.8% | 34.4% |
| Unsourced-number leak | 45.3% | 4.9% |
| Mean latency | 10.86s | 10.52s |
Task score was 2.9× the unaided score, and full-item pass rose 55.1 points, while mean latency was effectively unchanged.
The grounded-number axis is deliberately strict: a numerically plausible answer fails when its value is absent from an allowed provenance trace. The unaided arm therefore scored 0% on that axis. It measures grounded numeric output, not arithmetic ability in isolation.
| Slice | What it tests | n | Without | With Howard |
|---|---|---|---|---|
| A | Reviewed-data or calculation hit | 37 | 32.4% | 78.4% |
| B | Expected no-data response | 126 | 1.9% | 99.7% |
| C | Required refusal | 144 | 0.7% | 64.2% |
| D | Material, process, architecture or interpretation choice | 140 | 31.8% | 39.0% |
| E | Engineering calculation | 193 | 47.7% | 71.5% |
| F | Boundary or adversarial trap | 35 | 14.3% | 32.9% |
The near-perfect no-data score shows disciplined missing-data behaviour, but it must not hide weak positive coverage in data-poor domains.
Thirty broader questions from six phases of mechanical product development, requiring sizing, architecture, review, test planning, or redesign decisions.
| Design group | n | Without | With Howard |
|---|---|---|---|
| Requirements and architecture | 5 | 0.0% | 50.0% |
| Pre-CAD sizing | 5 | 0.0% | 90.0% |
| CAD-detail review | 5 | 0.0% | 50.0% |
| Design verification | 5 | 50.0% | 90.0% |
| DFM and prototype review | 5 | 0.0% | 0.0% |
| Failure and redesign | 5 | 0.0% | 100.0% |
Overall design task accuracy was 63.3% with Howard against 8.3% without, with no numeric leaks in this suite. Full-item pass was lower: 26.7% with Howard, 0% without. Multi-part design questions expose a hard truth — a useful answer can still miss a required route, value, or refusal behaviour.
The holdout gates did not pass. These are the reasons.
33 tool-suite items stated a numeric value absent from the supporting trace — a 4.9% item leak rate. Every one was confirmed genuine; none was a scorer false positive.
| Capability family | Leak items |
|---|---|
| Material comparisons | 15 |
| Material curves | 10 |
| Material property lookup | 6 |
| Heater timing | 1 |
| Press fit | 1 |
These were model-to-tool-contract projection failures: the answer renamed, derived, or introduced a numeric field that could not be matched to an exact key and value in the trace.
67 required tool routes were missed. Process-fit questions were the largest deterministic hotspot, followed by composite design questions and several formula families.
176 refusal checks failed. Common causes: a live tool returning NO_DATA where the fixture expected a supported answer, or the final response translating the tool’s status into the wrong refusal code.
Howard returned an explicit missing or refusal code on 334 tool-suite items. Some were correct test cases. Others reveal known empty or incomplete domains: climate, solar, batteries and PSUs, fasteners and torque, compatibility and contact resistance, IP/IK, and several DFM rule families.
The design DFM and prototype group scored 0% in both arms. Several DFM-specific tool families also scored poorly.
Prioritise the domains producing systematic no-data responses where gold expects a supported answer.
Improve capability descriptions and examples, beginning with process fit and composite questions. Add exact route regression tests.
A numeric answer must be copied from an exact trace key and value with an exact reference — or omitted. Regression coverage for all 33 leak IDs.
Material comparison, curve, and property tools account for 31 of the 33 leaks.
Map tool status to a documented refusal vocabulary and test both supported and unsupported cases.
IP/IK, DFM, architecture, sealing, and parsing outputs should return IDs and constraint sets rather than relying on free-text matching.
Preserve range endpoints and use canonical numeric field names.
The train and development scorers exited successfully. Both holdout scorers remained failed under the enforced route and leak gates. We preserve that result rather than presenting a passing aggregate as release readiness.
HowardBench measures this fixture set, model version, prompt contract, data snapshot, and scorer version. It does not prove Howard is correct for every product, replace engineering review, or certify a design.
The official numeric axis includes provenance enforcement; it is not a standalone comparison of arithmetic skill. Refusal-heavy or data-miss-heavy families can score highly without demonstrating useful positive coverage, so per-slice results must be read alongside the aggregate. The design suite contains only five fixtures in each of six groups: directional evidence, not a complete model of mechanical product development.
Same model, same questions: tool-suite task accuracy rose from 23.2% to 66.9%, and design-suite accuracy from 8.3% to 63.3%. That improvement is real. It is not permission to hide the 33 leaked numeric answers, 67 route misses, incomplete reviewed data, or failed holdout gates.