Notes
Working notes beside the benchmark: an introduction, and earlier essays on what AI work costs and how it is checked.
- What this benchmark measures, and why
Margin of Intelligence grades AI systems on finance work a professional would have to sign: planted errors in synthetic models and funding decisions, graded blind, published with the method and the evidence.
- Token Price Is Not Task Cost
A model's posted rate is only one input into the cost of useful AI work. The decision unit is the accepted task, after retries, tools, verification, review, and delay.
- AI Generates; Systems Verify; Humans Decide
A risk-tiered operating doctrine for using AI at scale: let models create options, make systems test what can be tested, and reserve consequential judgment for accountable people.
- The Exception Report Is the Product
In AI-native finance, value often comes from turning a complete population into a controlled queue of material exceptions with evidence, owners, and feedback—not from producing an autonomous answer.