Cricket Manager
IPL franchise simulation

The 0.98 was the problem.

I’m building a detailed simulation of running an IPL franchise, playable live with friends. A suspiciously clean model result was the clue that the game still needed work.


In short

Problem
How do you make a simulated cricketer behave like the real one — and how would you know if you had failed?
Constraint
T20 is the most chaotic format in the sport, and the engine has to be perfectly deterministic.
Decided
Adapt the rating model from the Earth Similarity Index. Abandon min-max normalisation for reference-value deviation. Make Spearman correlation the acceptance criterion rather than eyeballing. Event-source the auction.
Rejected
Keeping the min-max version because it was already built and its outputs looked reasonable — and reporting ρ = 0.98 as the headline validation result.
Hardest
Proving single-writer correctness before trusting the event log, and getting net run rate exactly right.
Outcome
2.07M real deliveries modelled. Direction proved, magnitude not. Nobody but me has played it.

Question

How do you make a simulated cricketer behave like the real one — and how would you know if you had failed?

The first half of that is a modelling problem and the second half is the harder one. Any management simulation can produce plausible-looking scorecards. Plausible is the trap: a scorecard that reads correctly tells you nothing about whether the engine underneath is reproducing the sport or merely imitating its output format.

So the project is really two things. A game — you own an IPL franchise, build a squad through a live auction, and watch a season play out ball by ball against other real owners. And a validation problem, which is where nearly all the interesting work went.

The reference points are Football Manager and OOTP Baseball rather than anything on a phone: for people who would rather agonise over whether to burn an RTM card on a death-overs specialist than bowl the yorker themselves.

Constraint

T20 is the most chaotic format in the sport, and the engine has to be perfectly deterministic.

Those pull in opposite directions, and holding both is the entire technical problem.

Determinism is non-negotiable for a multiplayer simulation: the same inputs must produce the same match everywhere, forever, or an auction result cannot be trusted and a bug cannot be reproduced. But the sport being modelled is one where a single over flips a match, and a system built for reproducibility will drift toward tidiness if nobody is watching the right number.

The second constraint is data volume, which is a gift with a bill attached. Roughly 797MB of ball-by-ball Cricsheet data — 11,190 match files across 13 T20 competitions, about 2.07 million individual deliveries — is enough to calibrate against reality rather than taste. It is also enough that "it looks about right" stops being an acceptable answer, because you now have the means to check.

Squad data covers all 10 IPL franchises and 248 active players with batting and bowling statistics each.

Decision

Ten decisions. The last one is the only one I would call judgment rather than engineering.

Decision 01

Steal the rating model from astronomy

Player ratings are built on an adaptation of the Earth Similarity Index — the astronomy metric that scores how closely an exoplanet resembles Earth on a 0-to-1 scale. Swap the planet for a cricketer and Earth for an ideal statistical profile, and you have a rating system with a defensible shape rather than a set of numbers someone felt were about right.

Batting Similarity Index and its bowling counterpart are the result. The appeal is not novelty; it is that the model forces you to state what an ideal player is, numerically and per format, before you can score anyone against it. A hand-tuned rating hides that judgment. This one has to publish it.

Decision 02

Abandon min-max normalisation for reference-value deviation

The first formulation was a min-max normalised weighted sum: each statistic scaled against its historical range, weighted, and combined across formats.

It works, and it has a flaw that matters. Min-max normalisation scores a player against the population, so a rating silently moves when the population changes, and a "90" means nothing more specific than "near the top of whoever else is in the dataset."

The refined model scores each statistic by how far it sits from a fixed ideal instead, so a rating means the same thing whoever else is in the game.

Traded away: automatic recalibration. A reference-value model has to be revisited by hand as the sport changes, where a min-max model drifts along with it for free.

Rejected: keeping the min-max version because it was already built and its outputs looked reasonable. Reasonable-looking output from a model whose scale has no fixed meaning is exactly the failure this project is about.

Decision 03

Publish the model's own weaknesses alongside it

The design documents state the limitations rather than leaving them to be discovered. The reference values and weights are subjective and high-leverage — small changes there move every rating in the system. The model collapses a multidimensional career into one number. It ignores opposition quality, conditions and era entirely. And it is most valid comparing players of similar role and vintage, which is a real restriction on a game whose auction puts a 2012 specialist next to a 2026 one.

Three extensions are specified and deliberately not built: a longevity factor for long careers, a peak-performance index derived from best-innings and best-match bowling figures, and a combined all-rounder index blending the batting and bowling scores.

7 more decisions inside Continue the Cricket Manager case study Decisions 04–10 · rating model · build phases · about 1,300 words Show all 7 remaining decisions ↓ Hide the full record ↑

Decision 04

Make Spearman correlation the acceptance criterion, not eyeballing

Calibration against the ball-by-ball corpus and real auction data produced three separate rating models — Ability, Market Value and Form — validated at Spearman correlations of 0.65 to 0.66 against real-world outcomes.

Splitting the rating into three is the substantive decision. A single number cannot simultaneously answer "how good is this player", "what will he cost" and "how is he playing right now", and collapsing them is how simulations end up with a squad of statistically excellent players who lose. Market value has to be able to disagree with ability — that gap is where the auction game actually lives.

0.65 is the honest number, not a flattering one. Checking ratings against real auction prices and real match outcomes means a 90-rated player has to have actually played like one, and that is a bar you can fail publicly.

Decision 05

Treat determinism as sacred in the engine

The ball-by-ball engine is fully deterministic, and the rules protecting that are absolute: one seeded source of randomness, nothing random outside it, and no I/O inside the engine at all.

Guarding it is a golden-master test suite pinned to one reproducible scorecard. Any change that shifts that scorecard is a behavioural change, whether or not it was meant to be one. Aggregate statistical output was validated to within ±0.2 percentage points of real-world tendencies.

Decision 06

Event-source the auction

The auction core is deterministic and event-sourced: bids are processed as an ordered event log rather than by mutating shared state. For a live room where several people bid within the same second, the log is the truth — replayable, auditable, and impossible to disagree about after the fact.

It carries the real mechanics rather than an abstraction of them: retentions before the open market, RTM cards, the capped, uncapped and overseas categories, and a hard budget cap so every rupee spent on one player is one you cannot spend on the next. Stress-tested with 8-bidder race scenarios.

Decision 07

Damp the AI bidders on purpose

AI franchises run a saturation damper built specifically to stop them hoarding. Without it, bots with a budget and an objective function buy everything worth buying, and the market stops behaving like a market.

The intended effect is that prices rise on genuine scarcity rather than because a bot got greedy — which matters because the auction is the marquee moment of the product. An unrealistic auction does not just play badly; it invalidates every squad decision that follows from it.

Decision 08

Get net run rate exactly right

The league engine uses circle-method round-robin scheduling and an IPL-style playoff bracket — Qualifier 1, Qualifier 2, Eliminator, Final. It also implements the net run rate rule that most implementations get wrong: a side bowled out counts as having faced its full quota of overs, not the number it actually used.

That single rule decides qualification in real seasons. Getting it wrong produces a league table that looks right all year and sends the wrong team to the playoffs.

Decision 09

Prove single-writer correctness before trusting the log

The engine and the auction were built against in-memory adapters first, so they could be tested without any real infrastructure. Wiring them to a real database and a real-time layer meant the event log's central guarantee had to be proved rather than assumed.

The database itself refuses two events at the same position in the same stream, and a test fires 8 simultaneous append attempts to prove it. If two bids can occupy the same version of the same stream, event sourcing has bought nothing.

Identity in the live room follows the same posture: server-derived identity — client-asserted user IDs are never trusted. In a room where identity determines who owns a bid, accepting the client's word for who it is would be the whole vulnerability.

Decision 10

Treat a suspiciously good score as a defect

The season-sanity probe compared simulated season outcomes against real-world plausibility by Spearman correlation and returned 0.885 to 0.981.

That is a very strong result, and it is the finding I acted on. In most simulations a 0.98 would close the ticket. In T20 it means the model is too orderly: a format famous for a single over flipping a match should not produce seasons that correlate that cleanly with expectation. Real T20 seasons contain results nobody would have ranked correctly beforehand.

So the open problem is named precisely: engine tilt-magnitude calibration. Earlier work proved the simulation is monotonic — better players do produce better outcomes, in the right direction. It never proved the magnitude of that tilt is realistic. Direction was verified; amplitude was not, and the season-sanity numbers are what turned that gap from a theoretical note into a concrete worry.

Rejected: reporting 0.98 as the headline validation result. It is the most flattering number the project has produced and it is evidence of a problem.

The rating model

The reference values are the argument. Everything else is arithmetic.

From dataset to acceptance criterion in the simulation's calibration chain A corpus of 2.07 million real T20 deliveries feeds a rating model built on reference-value deviation, adapted from the Earth Similarity Index after min-max normalisation was abandoned. The model drives a deterministic ball-by-ball engine, whose output is judged against a Spearman rank correlation threshold agreed in advance rather than by inspection. The resulting 0.98 correlation was treated as evidence of a defect rather than as a validation result. 2.07Mreal T20 deliveriesRATING MODELreference-value deviationadapted from the Earth Similarity Index;min-max normalisation abandonedENGINEdeterministic, ball-by-ballSPEARMAN ρacceptance criterionnot eyeballing — a threshold agreedbefore the result was knownρ = 0.98 came back. It was read as evidence of a leak rather than as a validation, and never published as the headline.Direction was proved. Magnitude was not.
Fig. 01 — The acceptance criterion was set before the number arrived. That is the only reason the number could be rejected.

The anchors are where the judgment lives, and they are chosen per format. That is what decides whether a model is a translation of the sport or an imposition on it: a T20 bowler is scored against what twenty overs make achievable, not against milestones borrowed from the longer formats, which would flatten every bowler in the game towards zero on that component.

What each phase built

Design documents precede implementation for anything touching the attribute-to-outcome mapping. That rule produced the order below.

Phase history and what each one proved
PhaseBuiltProved
RatingsBatting and bowling indices calibrated against the ball-by-ball corpus and real auction data; Ability, Market Value and Form modelsSpearman 0.65–0.66 vs real outcomes
Match engineDeterministic ball-by-ball simulationMonotonicity; aggregates within ±0.2pp; a pinned golden master
AuctionEvent-sourced auction with retentions, RTM and damped AI biddingCorrect under 8-bidder race conditions
LeagueRound-robin scheduling, exact NRR, IPL playoff bracketSeason sanity at Spearman 0.885–0.981 — and the flag it raised
InfrastructureModular design with in-memory adaptersTestability without real infrastructure
PersistenceA real database and real-time layer behind the same interfacesSingle-writer correctness under 8 concurrent appends
Auth & roomsAuthenticated live auction roomServer-derived identity end to end
StatsLeaderboards and value-for-money analytics on the same career data as the ratingsPoints per crore as an auction-strategy measure

How it is built

A TypeScript monorepo, with a real-time layer for the live auction room.

Development runs through Claude Code against a persistent context file holding stable project information, with task-scoped prompts issued per phase — a deliberate discipline for keeping long-running, multi-session AI-assisted development coherent rather than letting each session re-derive the architecture. Every phase closes with a summary and an open-items review before the next one opens. The tilt-magnitude problem stayed visible across three phases precisely because that review exists.

What it looks like

The interface is a control room rather than an arcade: dark, data-dense, purposeful, closer in spirit to Cricket Captain than to anything designed for a thirty-second session. It is built to be lived in for a whole season, and the visual design follows from that rather than from what screenshots well.

Outcome

  • 2.07MDeliveries calibrated against
  • 11,190Match files, 13 competitions
  • 0.65Spearman, ratings vs real outcomes
  • ±0.2ppAggregate output vs reality
  • 248Players, 10 franchises
  • 8Concurrent appends, log correct

Every system described above was built and tested: the rating models against real outcomes, the engine against real aggregates, the auction under race conditions, the league against real qualification rules, the event log under concurrency, and the auction room under an authenticated multiplayer session.

What is not settled is the one thing the strongest number exposed. Tilt magnitude remains open, and it is the current focus of refinement rather than a missing feature — the seasons are too well-behaved for the sport they model.

Nobody but me has played it. The build evidence here is unusually complete and the usage evidence is nil — no seasons completed outside testing, no live auction room with other people in it. Whether it becomes something other people play is still open.

Why a simulation is worth building at all

This is the least commercial thing on this site and the most useful thing I have built for thinking. A simulation is systems design with nowhere to hide: you are not shipping features, you are building a small economy with its own internal logic, and if any part of it is unbalanced players find the exploit within hours.

Every subsystem turns out to be a familiar problem wearing a cricket shirt. The auction is a market. The rating system is a model, with the same bias and edge-case failures as any other model. Scheduling is constraint satisfaction. The domain changes; the thinking does not — which is the argument for building outside your field rather than a hobby that needs excusing.

Direction was proved. Magnitude was not. A simulation that gets the ranking right and the chaos wrong is a spreadsheet wearing a scorecard.

typescript · next.js · real-time multiplayer · event sourcing · deterministic simulation · spearman validation · cricsheet

Next