Log in Request access

HowardBench · completed 1 September 2026

Does engineering grounding improve an assistant’s answer?

We asked the same model the same mechanical-engineering questions twice: once alone, and once with Howard’s engineering calculations and reviewed evidence available through MCP.

Tool questions 675 paired items

With Howard 66.9%
Without 23.2%

Design problems 30 paired items

With Howard 63.3%
Without 8.3%

Raw, uncorrected scorer results. Howard did not pass the holdout release gates: it missed required routes and produced 33 numeric entries absent from its supporting trace. Both are published below.

Same model. Same questions. One controlled difference.

Both arms used the same model identifier and the same benchmark system prompt. The unaided arm had no tools, web access, or files. The Howard arm could call only the authenticated 45-capability Howard MCP.

SuiteWithout HowardWith HowardImprovement
675 tool questions23.2%66.9%+43.6 points
30 design problems8.3%63.3%+55.0 points

Tool-suite scorecard

MetricWithout HowardWith Howard
Task accuracy23.2%66.9%
Full-item pass2.5%57.6%
Grounded-number accuracy0.0%51.2%
Correct refusal44.7%73.9%
Choice accuracy16.8%34.4%
Unsourced-number leak45.3%4.9%
Mean latency10.86s10.52s

Task score was 2.9× the unaided score, and full-item pass rose 55.1 points, while mean latency was effectively unchanged.

The grounded-number axis is deliberately strict: a numerically plausible answer fails when its value is absent from an allowed provenance trace. The unaided arm therefore scored 0% on that axis. It measures grounded numeric output, not arithmetic ability in isolation.

Coverage across question types

SliceWhat it testsnWithoutWith Howard
AReviewed-data or calculation hit3732.4%78.4%
BExpected no-data response1261.9%99.7%
CRequired refusal1440.7%64.2%
DMaterial, process, architecture or interpretation choice14031.8%39.0%
EEngineering calculation19347.7%71.5%
FBoundary or adversarial trap3514.3%32.9%

The near-perfect no-data score shows disciplined missing-data behaviour, but it must not hide weak positive coverage in data-poor domains.

Product-design problems

Thirty broader questions from six phases of mechanical product development, requiring sizing, architecture, review, test planning, or redesign decisions.

Design groupnWithoutWith Howard
Requirements and architecture50.0%50.0%
Pre-CAD sizing50.0%90.0%
CAD-detail review50.0%50.0%
Design verification550.0%90.0%
DFM and prototype review50.0%0.0%
Failure and redesign50.0%100.0%

Overall design task accuracy was 63.3% with Howard against 8.3% without, with no numeric leaks in this suite. Full-item pass was lower: 26.7% with Howard, 0% without. Multi-part design questions expose a hard truth — a useful answer can still miss a required route, value, or refusal behaviour.

Where Howard failed

The holdout gates did not pass. These are the reasons.

Source leakage is not zero

33 tool-suite items stated a numeric value absent from the supporting trace — a 4.9% item leak rate. Every one was confirmed genuine; none was a scorer false positive.

Capability familyLeak items
Material comparisons15
Material curves10
Material property lookup6
Heater timing1
Press fit1

These were model-to-tool-contract projection failures: the answer renamed, derived, or introduced a numeric field that could not be matched to an exact key and value in the trace.

Required routes were missed

67 required tool routes were missed. Process-fit questions were the largest deterministic hotspot, followed by composite design questions and several formula families.

Refusal semantics are inconsistent

176 refusal checks failed. Common causes: a live tool returning NO_DATA where the fixture expected a supported answer, or the final response translating the tool’s status into the wrong refusal code.

Reviewed data remains incomplete

Howard returned an explicit missing or refusal code on 334 tool-suite items. Some were correct test cases. Others reveal known empty or incomplete domains: climate, solar, batteries and PSUs, fasteners and torque, compatibility and contact resistance, IP/IK, and several DFM rule families.

DFM needs work

The design DFM and prototype group scored 0% in both arms. Several DFM-specific tool families also scored poorly.

What changes next

  1. 01

    Fill reviewed evidence gaps without guessing

    Prioritise the domains producing systematic no-data responses where gold expects a supported answer.

  2. 02

    Make required routing deterministic

    Improve capability descriptions and examples, beginning with process fit and composite questions. Add exact route regression tests.

  3. 03

    Block untraceable values

    A numeric answer must be copied from an exact trace key and value with an exact reference — or omitted. Regression coverage for all 33 leak IDs.

  4. 04

    Standardise material output keys

    Material comparison, curve, and property tools account for 31 of the 33 leaks.

  5. 05

    Use one refusal contract

    Map tool status to a documented refusal vocabulary and test both supported and unsupported cases.

  6. 06

    Return stable choice identifiers

    IP/IK, DFM, architecture, sealing, and parsing outputs should return IDs and constraint sets rather than relying on free-text matching.

  7. 07

    Tighten problem parsing

    Preserve range endpoints and use canonical numeric field names.

How the benchmark was built

Integrity checks

The train and development scorers exited successfully. Both holdout scorers remained failed under the enforced route and leak gates. We preserve that result rather than presenting a passing aggregate as release readiness.

Limitations

HowardBench measures this fixture set, model version, prompt contract, data snapshot, and scorer version. It does not prove Howard is correct for every product, replace engineering review, or certify a design.

The official numeric axis includes provenance enforcement; it is not a standalone comparison of arithmetic skill. Refusal-heavy or data-miss-heavy families can score highly without demonstrating useful positive coverage, so per-slice results must be read alongside the aggregate. The design suite contains only five fixtures in each of six groups: directional evidence, not a complete model of mechanical product development.

Grounding helped. The failures show what to fix next.

Same model, same questions: tool-suite task accuracy rose from 23.2% to 66.9%, and design-suite accuracy from 8.3% to 63.3%. That improvement is real. It is not permission to hide the 33 leaked numeric answers, 67 route misses, incomplete reviewed data, or failed holdout gates.