Introduction

What this benchmark measures, and why

Margin of Intelligence grades AI systems on finance work a professional would have to sign: planted errors in synthetic models and funding decisions, graded blind, published with the method and the evidence.

By Mohammed Abdul GaffarPublished Updated

I review and build financial models for a living. The question I am asked most often now is whether AI can do this work. The answers I hear are demonstrations: a model built in minutes, a review that reads well. A demonstration shows what a system can do on a good day. It does not show what it misses, what it makes up, what it costs, or how much checking is left for the person who has to sign.

Margin of Intelligence exists to measure those four things.

The method is plain. Take a synthetic company. Build its three-statement model and its valuation. Plant a small number of realistic errors: a tax charge on the wrong base, a terminal growth rate the discount rate cannot support, an interest line typed in where a link to the debt schedule should be. Give the same file to several AI systems under the same instructions, three times each. Grade what comes back, blind, against a sealed key. Count what was caught, what was missed and how much it mattered, what was invented, what it cost, and how many minutes of review the output still needs. Publish the method, the cases, the evidence and the limitations with every result.

Two things matter as much as the count. Whether a system raises false alarms on a clean model, because a reviewer who cries wolf costs an afternoon. And whether it gives the same answer twice, because a tool that cannot repeat itself cannot sit inside a process.

Everything here is synthetic. No employer, client or bank document appears on this site, and none ever will. That is not only a matter of confidentiality. A synthetic model can carry exactly the errors we choose, and that is what makes grading possible.

The first study is in development: two synthetic companies and ten funding decisions, across the systems people are actually using. First results are planned for October 2026. Until then, the method describes what will be measured and what will not, and there are no numbers on this site, because there are no results yet.

The essays beside this note were written in August 2026, before the benchmark took this shape. They are about what AI work costs after retries and review, how verification should be organised, and why finance automation ends up as an exception queue. They stay here because they are the reasoning behind the measures.

If you review models for a living, tell me which error you would plant.

Corrections

No corrections recorded.

Report a correction via LinkedIn with the note title and the supporting source.