We Built an AI Federal Reserve

Jeevan Sandhu and Haardik Garg22 July 2026
The Marriner S. Eccles Federal Reserve Board Building, Washington, D.C.

Overview

Eight times a year, a committee at the Federal Reserve takes inflation, employment, growth, market pricing, and a stack of narrative evidence that rarely points the same way, and compresses all of it into one number: the change in the federal funds target rate. That number reprices mortgages, currencies, and the cost of government debt across the world economy. It is reached by people who weigh the risks differently from one another, behind a closed door, and explained to the public only after the fact.

We wanted to know how much of that decision could be rebuilt from the outside, using only what was public and knowable in advance. So we built OpenFed and ran it against the entire modern record: 263 FOMC meetings from February 1994 through June 2026. For each one, the system saw only the evidence available a week before the meeting, with the real outcome held back, and had to reach the call on its own.

It matched the Fed’s actual decision exactly 259 times out of 263. That is 98.48 percent, spanning four chairs, two financial crises, and a pandemic. But the reconstruction was only the audition. What OpenFed was built to do starts today.

It Goes Live This Week

Starting today, it does exactly what it did across those 263 historical meetings, except forward and in full view. Seven days before each scheduled FOMC decision, the committee convenes, deliberates, and publishes everything: the five archetypes arguing, both rounds of votes, its comparison against a transparent policy rule, and the chair’s statement on the decision. It reaches its call a week ahead of the Fed, every time, with nothing held back.

The first live run is today, Wednesday, July 22, at 2:00 PM ET, seven days ahead of the Fed’s July decision. Everything the system produces will be posted as it happens.

Then we are doing something that has not been tried before. On Tuesday, July 28 at 3:00 PM ET, the day before the FOMC announces, OpenFed’s AI chair will hold a live meeting. The chair will deliver its statement on the July rate decision and then take questions directly from the room, over Zoom. Anyone can attend and ask: economists, policymakers, journalists, students, and citizens who simply want to press an autonomous system on a decision that reaches their mortgage and their paycheck. You ask the question, the AI chair will answer, in real time.

Registration for the meeting is open at luma.com/pbex3q9z.

How It Works

OpenFed runs two engines, and they never see each other’s work.

The first is a committee of five fixed policy archetypes: an inflation hawk, a labor-market dove, a financial-stability voice, a data-dependent centrist, and a market-signal realist. They are perspectives rather than impersonations of real officials, and they read the same meeting packet before weighing it through their own priorities and casting an independent vote. Then they read one another’s reasoning and vote a second time, after which a plurality rule settles the decision. Every vote, every argument, and every dissent is kept.

The packet holds what a committee member could plausibly have known seven days out: core inflation, unemployment, the natural-rate estimate, financial-stress indicators, the yield curve, credit spreads, market-implied rate pressure, the latest Beige Book, and a summary of the previous meeting.

The second engine sets the Fed aside entirely. It is deterministic, with no personas and no debate, and it answers a different question: given a transparent dual-mandate objective fixed in advance, which action minimizes the loss? It balances the projected inflation gap, the projected employment gap, the distance from a neutral rate, a penalty for tightening into financial stress, and a small penalty for moving too abruptly. The result is fully auditable arithmetic.

What It Got Right

Begin with the misses, since there are only four, and none of them is off by more than a single quarter-point.

Twice in 1994, OpenFed expected a quarter-point hike where the Fed chose to hold. In January 1995 it called a quarter-point tightening when the Fed moved a half, and in September 2024 it called a quarter-point cut when the Fed cut a half. That is the complete error set for the modern era. Averaged across all 263 meetings, the typical miss is 0.38 basis points, and the direction of the move is right 261 times.

What is striking is where the errors are absent. OpenFed was exact through the 2008 collapse, through the years pinned at the zero lower bound, through the pandemic response, and through the whole 2022 hiking cycle, the sharpest tightening in four decades. The meetings that felt like the ground was moving turned out to carry a clear forcing logic that the indicators expressed on their own. Three of the four misses instead fall in the quiet middle of the 1990s, where the data was mild and a quarter-point could reasonably have gone either way.

For scale, the strongest prior systems in this line of work evaluated 8, 16, and 60 meetings. OpenFed runs all 263 and keeps every individual vote, and on every window where a direct comparison is possible, it comes out on top. MiniFed matched 6 of 8 real 2018 decisions; OpenFed matches all 8. FedSight AI’s best-performing variant hit 15 of 16 meetings from 2023 to 2024; OpenFed matches that mark exactly, and is correct on the direction of all 16. Takano et al. reported a macro F1 of 0.476 across their 60-meeting sample; OpenFed is exact on 100 percent of meetings across that same stretch of the record. No published system in this line of work has matched a longer stretch of FOMC history, and none has matched it this precisely. It also leaves the naive baselines far behind: always guessing “hold” is correct 67 percent of the time, because the Fed usually does nothing, and the distance between that baseline and OpenFed is about as statistically decisive as a result gets.

What the Arguing Was Worth

We added the second voting round because the research consensus holds that debate sharpens reasoning and lets models catch one another’s mistakes. So we let them argue, and then we measured whether it helped.

Out of 1,315 individual votes, the debate changed 35. At the committee level, once everything had reshuffled, the second round improved the final answer on a net of essentially zero meetings. Whatever the committee knew, it already knew from the first independent vote, and the discussion mostly moved positions sideways.

The more useful discovery sits right next to that one. When the five archetypes agreed, they were almost never wrong: a single error across 211 unanimous meetings. When they split, the error rate rose more than twelvefold. The disagreement did not repair the decision, but it reliably marked the meetings worth distrusting.

The same robustness shows up when we remove voices. Pulling any one of the five archetypes out and re-running the aggregation leaves the decision unchanged in 249 of 263 meetings. No single perspective is load-bearing. The hawk carries the most weight, and even removing the hawk flips only ten calls out of 263. The committee’s judgment is spread across the group rather than captured by any one voice, which is exactly the property you would want from a body meant to hold a balance of risks.

The Wedge

The committee reconstructs what the Fed did, with 98 percent fidelity. The deterministic rule computes what a written-down objective wants, indifferent to what the Fed actually chose. Set the two side by side and you can read the difference between them at every meeting across 32 years, with a direction and a magnitude in basis points. We call it the wedge.

The two disagree in 156 of the 263 meetings, and the disagreement leans one way. The reconstructed Fed comes out more hawkish than the rule far more often than the reverse, its real choices sitting tighter than the transparent objective would prescribe. The pattern shifts with the era. During the years at the zero lower bound the rule and the realized action agree almost perfectly, because pinned at zero there is little room to differ. Through the 2016 to 2019 normalization they come apart sharply, with the rule matching the Fed’s actual move only 5 times in 32 meetings. And the divergence is systematic: regress the wedge on the state of the economy and it tracks the market-implied rate expectations that the committee’s realist responds to and the deterministic objective simply does not contain. The reconstructed Fed is listening to what markets are pricing. The written rule is not.

None of this says the Fed was wrong. The rule is one explicit objective among many defensible ones, and a central bank’s mandate does not reduce to five coefficients, which is exactly why the separation is worth having. The wedge is not a scorecard against the truth; it is a measurement of the distance between an institution’s revealed behavior and one stated objective, and it holds up under pressure. Perturb all five weights of the loss function by 25 percent in every direction at once, and 162 of the recommendations do not move at all, while none of the rest moves by more than a single quarter-point. The gap is a property of the decisions, not of the knobs we chose. A forecast can tell you what the Fed will do. The wedge tells you how far the institution’s judgment sits from its own stated objective, meeting by meeting.

Conclusion

OpenFed rebuilt 32 years of Federal Reserve decisions from five perspectives and a week-old information packet, and matched all but four of them within a single quarter-point. Along the way it showed that deliberation barely changed the answer yet flagged the hardest calls, that no single voice on the committee was indispensable, and that the institution’s real choices run systematically tighter than a transparent objective says they should.

It is live now at openfed.app, the full paper is at files.openfed.app/paper.pdf, and the AI chair takes your questions on July 28. Register at luma.com/pbex3q9z.