# Simulating Markets Instead of Fitting Curves to Them

> MiroFish replaces curve-fitting with agent-based social simulation for demand forecasting, seeded with FRED data, macro signals, and decades of history. What the architecture gets right, what it still cannot do, and where the evidence actually holds up.

- Published: 2025-05-09
- Reading time: 17 min read
- Tags: ML, Engineering
- Author: Caio Theodoro (https://caio.theodoro.dev/about.md)
- Canonical HTML: https://caio.theodoro.dev/blog/mirofish-demand-forecasting

---

## How MiroFish uses swarm simulation for demand forecasting with FRED, macro signals, and 40 years of history


---


## The model that works until it doesn't


Most demand forecasting models are embarrassingly bad at the moments when accuracy matters most, and the industry mostly knows this and mostly looks away.


The standard toolkit does its job for 90% of the distribution: LightGBM catches seasonal patterns, LSTMs stacked on transformers track the trend, gradient-boosted trees absorb whatever FRED macro series you feed them. In quarterly reviews, the MAPE looks fine. Stakeholders nod. Then a tariff announcement hits, a Fed decision comes in sideways, or a TikTok goes viral and shifts consumer sentiment overnight. The model that was nailing 4.2% last month is off by 35%. You're staring at a warehouse full of inventory nobody wants, or an empty shelf that should have been restocked three weeks ago.


The problem is not implementation quality. The problem is the framing. Traditional demand forecasting treats the world like a physics problem: deterministic, reducible to equations, solvable with enough inputs. But demand is driven by people who react to each other, who form opinions based on what their neighbors are buying, who panic-buy when the news cycle gets scary. No amount of feature engineering accounts for the social dynamics of a disruption. The cascade, importers scrambling to pre-order, retailers adjusting shelf space, consumers seeing news coverage and making individual decisions that collectively reshape the market, is not visible in the historical correlation.


MiroFish addresses that gap by simulating the people instead of fitting a curve to what they did last time.


---


## Every major forecasting approach shares one assumption


Every major forecasting approach of the last thirty years has shared the same premise: the future is, at least partially, legible from the past. MiroFish quietly abandons that premise. Rather than fitting a model to historical data, it asks what would happen if you simulated the people.


MiroFish is an open-source swarm intelligence engine built by Guo Hangjiang, an undergraduate at Beijing University of Posts and Telecommunications, in roughly ten days. Within 24 hours of a rough demo, Chen Tianqiao, the billionaire founder of Shanda Group, committed around $4.1 million to incubate the project. It hit #1 on GitHub's global trending list above OpenAI and Google repositories and has accumulated over 45,000 stars as of this writing.


The core mechanism: you provide "seed material," meaning news articles, financial reports, policy documents, datasets, market research, or whatever describes the current state of the world relevant to your prediction question. You describe the question in natural language. From there, the engine uses [GraphRAG](https://arxiv.org/abs/2404.16130), a retrieval approach that builds a structured knowledge graph rather than searching a flat index, to parse your input and extract entities, relationships, and pressures. That graph becomes the world model from which agent personas are generated, each with distinct backgrounds, stances, and behavioral logic. An Environment Configuration Agent sets the simulation parameters and rules.


Those agents then populate two parallel simulated platforms simultaneously, one shaped like Twitter's short-form dynamics and one like Reddit's threaded discussions, where they post, debate, argue, form coalitions, shift opinions, and influence each other. Their memories update continuously as simulated time advances via Zep Cloud, which means a simulated analyst who read a bearish macro report in round 3 still carries that read when they encounter a competitor announcement in round 47. A dedicated ReportAgent watches the whole thing unfold and synthesizes what the population collectively arrived at. And throughout, the simulation remains open: you can enter it, interview any individual agent, inject new variables, ask the ReportAgent follow-up questions, and watch what happens.


The simulation engine underneath is OASIS (Open Agent Social Interaction Simulations) from CAMEL-AI, which supports up to one million agents with 23 different social actions, including following, commenting, reposting, liking, muting, and searching.


---


## Why emergence is the point


Traditional demand forecasting fails at inflection points for an architectural reason: statistical models are trained on historical correlations, and correlations break down when the underlying causal dynamics shift. The model has never seen this particular combination of macro conditions, competitive pressure, and sentiment cascade, because it has never happened before.


Consider a practical example. You're forecasting demand for consumer electronics ahead of the holiday season. Your model has CPI, consumer sentiment index, unemployment rate, and three years of historical sales data. Then, a week before Black Friday, the Fed announces an unexpected rate hold, a major competitor launches a surprise price war, and a prominent tech reviewer publishes a scathing video that goes viral. Your ARIMA model doesn't know what to do with any of this. Your gradient-boosted tree might catch the macro shift a quarter later, when the CPI data finally reflects it. By then, the inventory decision is already made.


MiroFish handles this differently because it's not fitting curves to historical data. It's simulating how different market participants react to each other as the scenario unfolds. You seed it with the current macro environment, the competitors, the social media sentiment, and you watch thousands of simulated consumers, retailers, analysts, and media outlets interact. What emerges isn't a point estimate. It's a distribution of plausible futures shaped by the same social behavior that drives demand.


Demand is an emergent property of collective human behavior. It is not produced by any single consumer or retailer or analyst, but by the interactions between all of them at once. Which means the honest way to forecast it is not to fit a curve to its past trajectory, but to simulate the system that generates it — swarm intelligence isn't a loose metaphor for that problem, it's a direct architectural fit.


---


## How macro data, FRED, and historical series become seed material


MiroFish isn't a plug-and-play forecasting API. It's an engine that requires thoughtful input design, and this is where the demand forecasting application becomes powerful: the quality of your seed material determines the quality of your simulation.


### FRED as world-building data, not model inputs


The Federal Reserve Economic Data (FRED) database contains over 800,000 time series covering everything from the federal funds rate to county-level unemployment to the University of Michigan Consumer Sentiment Index. Most ML practitioners dump a handful of FRED series into their feature matrix and call it a day. MiroFish inverts this relationship.


Instead of treating FRED data as numerical inputs in a regression, you use it as context for world-building. You're not feeding the Consumer Price Index into a model as a floating-point number. You're writing the seed material that says: "CPI has risen 3.8% year-over-year, the highest in six months. Wages have stagnated after inflation. Grocery prices are up 5.2% while energy costs remain elevated. The Fed has held rates steady at 4.25–4.50% through the first quarter of 2026, and forward guidance suggests two cuts are likely before year-end."


This matters because the agents respond to this information the way people do: not as coefficients in a regression equation, but as contextual knowledge that shapes their behavior, risk tolerance, and purchasing decisions. A simulated retail consumer who knows that inflation is sticky and wages are flat behaves very differently from one who knows that prices are falling and employment is strong. And when thousands of these agents interact, the collective demand signal that emerges is grounded in behavioral responses to actual economic conditions.


The FRED series I've found most useful for demand forecasting seed material fall into four categories. Consumer-side indicators, CPIAUCSL, PCE, DSPIC96, UMCSENT, and RSXFS, tell your agents what it feels like to be a consumer right now. Labor market indicators, including UNRATE, PAYEMS, JTSJOL, and JTSQUR, drive the consumer confidence that shapes spending behavior; some studies have achieved forecasting error reductions of over 50% by incorporating leading macroeconomic indicators of this kind. Financial conditions indicators, such as FEDFUNDS, DGS10, SP500, and the Financial Conditions Index, shape business investment, credit availability, and the mood of institutional players. Supply-side indicators, including PPIACO, INDPRO, ISRATIO, and import price indexes, set the constraints on what's available, at what cost.


The trick is not to dump all of these into a single seed document, but to synthesize them into a coherent narrative of the current economic environment. MiroFish's GraphRAG will do some of this extraction, but the more structured and narrative your seed material is, the richer the knowledge graph becomes, and the more realistic the agent behavior.


### Historical demand patterns as agent memory


MiroFish agents have persistent memory via Zep Cloud, and you can seed that memory with historical patterns. This is where the demand forecasting application starts to matter.


Rather than feeding a time series into a statistical model, you embed historical demand patterns into the seed material as context that agents carry. For instance: "This product category has historically seen a 22% demand spike in Q4, followed by a 15% pullback in Q1. During the 2023 holiday season, demand exceeded forecasts by 31% because of a viral social media campaign. The 2024 season was flat as post-pandemic normalization continued."


The agents don't run a statistical model on this history. They carry it as contextual knowledge, the same way an experienced retail buyer or supply chain manager carries years of pattern recognition in their head. When the simulation runs, their decisions are informed by this history but not mechanically determined by it. They can deviate when the simulated conditions warrant it, the same way an experienced human forecaster overrides the model when they sense something changing.


### Dynamic variable injection for scenario testing


One of the most useful capabilities for demand forecasting is what the project describes as the "God's-eye view," the ability to dynamically inject variables into a running simulation. This maps directly to scenario planning and stress testing.


Say your simulation is running a baseline demand forecast for Q3. Midway through, you inject: "Breaking: Major port strike shuts down West Coast shipping for two weeks." The agents, importers, retailers, consumers, and media, react as the event enters the simulation. Supply chain anxiety cascades through the system: some consumers panic-buy, some retailers switch suppliers, some delay purchases entirely. The emergent demand pattern shifts in ways that no static model could anticipate, because the cascade effects are driven by agent-to-agent interaction.


You can pull these injection scenarios from anywhere: breaking news feeds, social media sentiment shifts, commodity price spikes, geopolitical developments, competitor announcements. What the injection mechanism tests is your demand forecast against exactly the kind of exogenous shocks that break traditional models.


---


## The architecture in practice


Having spent a few weeks prototyping with MiroFish, this is the pipeline architecture I've converged on for a demand forecasting application. MiroFish is at v0.1.2, so this isn't a production-ready blueprint, but the structure is sound.


The foundation is data ingestion and synthesis: pull the latest FRED data via the FRED API, pull internal sales history from your data warehouse, pull social sentiment from your monitoring tools, pull competitor intelligence from your market research feeds, and synthesize all of it into a structured seed document. That document is part narrative, part data, part knowledge graph primer. This is the state of the world your simulation will be built on, and it is the single most important artifact in the entire pipeline. A poorly constructed world produces agents whose behavior doesn't map to reality, and unlike a statistical model, there is no historical anchor to catch the error.


From there, simulation configuration: define your agent population. For a consumer goods demand forecast, a useful population might include a mix of consumer archetypes, such as price-sensitive, brand-loyal, impulse-driven, and research-heavy, alongside a cohort of retail decision-makers, a handful of market analysts, some media agents who amplify signals, and institutional buyers or wholesale agents. The GraphRAG step helps generate these personas, but reviewing and tuning them to reflect your actual market is where the work is.


Then the simulation runs. Agents interact, debate, form opinions, change their minds. The temporal memory updates continuously. No single agent is predicting demand; demand emerges from the collective behavior of agents making individual decisions, exactly like markets — you're growing a market rather than fitting a model to one.


The analysis step is where MiroFish offers something different from a statistical approach: you can enter the simulation and interview specific agents to understand why the demand forecast came out the way it did. Ask the simulated price-sensitive consumer why they delayed their purchase. Ask the retail buyer why they over-ordered. The explanatory power here exceeds what SHAP values on a gradient-boosted tree can give you.


The final layer is scenario branching: run the simulation multiple times with different injected variables. Each branch produces a scenario forecast, and the distribution across scenarios gives you something traditional forecasting rarely provides, an uncertainty estimate grounded in behavioral dynamics rather than statistical confidence intervals.


```mermaid
flowchart TD
  A([Seed Material]) --> B[GraphRAG]
  B --> C[World Model]
  C --> D[Agent Personas]
  D --> E{Simulation}
  E --> F[Twitter Platform]
  E --> G[Reddit Platform]
  F --> H[ReportAgent]
  G --> H
  H --> I([Forecast & Scenarios])
```


---


## The Polymarket evidence


The most compelling external validation of MiroFish's forecasting architecture comes from prediction markets, which is a useful proving ground because prediction markets are also crowd-behavior systems, and MiroFish is a crowd-behavior simulator.


Multiple developers have connected MiroFish to Polymarket trading bots. One well-documented case involved a trader who simulated 2,847 digital agents before every trade on Polymarket's short-term Bitcoin price markets. Over 338 trades, they reported $4,266 in profit, with one position returning 1,655% in five minutes. The edge was conceptually simple: when the simulated crowd's consensus diverged from what Polymarket was pricing, the bot entered the trade.


A more rigorous experiment involved simulating 200 agents to predict whether maritime shipping in the Strait of Hormuz would return to normal by the end of April 2026. The researcher seeded 200 agents with different roles, including government officials, media, military, energy companies, traders, and citizens, and ran 100+ rounds of interaction. The organic group consensus was 47.9% probability. Polymarket's market price was 31%, a 16.9 percentage point gap. The most cautious agents, the pessimists whose organic expressions most closely matched Polymarket's pricing, were the ones with the most domain expertise who participated naturally rather than being interviewed. The simulation didn't just produce a number; it produced a legible sociology of the disagreement.


This isn't a proof of concept for financial trading. It is evidence that swarm intelligence can produce useful signals in markets driven by human behavior, which is precisely the mechanism at work in demand forecasting.


---


## Where the architecture has an edge


The advantages show up most clearly at cascade events: tariff announcements, surprise Fed decisions, viral social moments that shift consumer sentiment in ways a survey index won't capture for weeks. A traditional model sees a feature value change in next quarter's data. A swarm simulation watches the cascade unfold agent by agent.


Sentiment contagion is the second place where the architecture has no statistical equivalent. Consumer sentiment isn't just a number in a survey. It spreads through social networks, gets amplified by media coverage, and creates feedback loops that can accelerate or dampen demand in ways that don't appear in any historical correlation until after they've already happened. A MiroFish simulation captures this contagion directly, through agents influencing agents.


The third structural advantage is heterogeneous agent behavior. Markets are not composed of a single representative consumer. They contain price-sensitive shoppers, brand loyalists, impulse buyers, and careful researchers, and these populations respond very differently to the same stimulus. A swarm simulation with diverse agent personas captures how the interactions between these groups produce the aggregate demand signal, in a way that treating consumer sentiment as a scalar feature in a regression simply cannot.


There is also what I'd call narrative-level explanation. When a traditional model's forecast is wrong, you can look at feature importances, but you can't ask it why. When a MiroFish simulation produces an unexpected result, you can go into the simulation, interview the agents, and understand the social dynamics that produced the outcome. That matters for supply chain decision-making, where trust in the forecast is as important as the forecast itself.


---


## What MiroFish cannot do yet


MiroFish, in its current state, has no published benchmarks against historical demand datasets. The Polymarket results are anecdotal. The system produces plausible scenarios, and those scenarios feel behaviorally realistic, but "feels right" is not validation. No one has yet run MiroFish against the 2020 pandemic demand shock or the 2021 supply chain crisis and measured how well the simulated emergent behavior matches what actually happened. Until that work exists, the architecture is promising, not proven.


The LLM cost structure is a constraint. The docs recommend starting with fewer than 40 simulation rounds for a reason: hundreds of LLM API calls per run. For enterprise-scale demand forecasting across thousands of SKUs on a daily cadence, the compute economics don't work yet. This will change as inference costs continue to fall, but it shapes where the approach is applicable today.


Seed material quality is everything, and there is no statistical fallback. A poorly constructed world produces confidently wrong simulations with no historical anchor to catch the error. This requires domain expertise in constructing the input, which means the approach is not easily operationalized by teams without that expertise.


And it is not a replacement for statistical models. MiroFish is strongest on the qualitative, behavioral, emergent dynamics that traditional models miss. It is weaker on the precise quantitative point estimates that supply chains need for daily inventory management. I think the most promising near-term approach is combining swarm simulation with conventional forecasting to cross-validate outputs: the statistical model provides the baseline point estimate, the swarm simulation provides the scenario distribution and the behavioral uncertainty band around it.


---


## What comes next


The convergence I'm watching is between swarm intelligence engines like MiroFish and the existing infrastructure of applied demand forecasting: FRED and macroeconomic data pipelines, alternative data feeds, and traditional forecasting stacks. I think this combination points to a different architecture for demand forecasting, not immediately and not cleanly, but directionally.


The specific developments that would move this from promising to practical are straightforward: automated FRED-to-seed-material pipelines that continuously synthesize the latest macroeconomic data into narrative context documents, hybrid ensemble architectures where a traditional model provides the baseline and MiroFish provides the scenario distribution and behavioral uncertainty band, and domain-specific agent libraries pre-configured for common demand forecasting contexts such as retail, CPG, automotive, pharma, and tech hardware.


The most important piece is calibration against historical shocks. Running MiroFish simulations against known disruptions, such as the 2020 pandemic, the 2021 supply chain crisis, the 2022 inflation spike, and the 2025 tariff waves, and measuring how well the simulated emergent behavior matches what actually happened is the path to the benchmarks the system currently lacks. That work is the difference between an interesting architecture and a trustworthy one.


---


## Where this leaves the field


Most engineering effort in demand forecasting goes into minimizing the failure mode at inflection points, not eliminating it: better inputs, tighter regularization, more training data, all still fit to a curve. The assumption that historical correlation is the right substrate for forecasting rarely gets questioned.


MiroFish questions it by replacing curve-fitting with simulation of the collective behavior that produces demand in the first place. The architectural insight holds up. What doesn't yet exist is the proof: MiroFish is at v0.1.2, there are no published benchmarks against historical demand shocks, and the cost structure doesn't support production-scale daily use. Until someone runs it against 2020 or 2021 and measures the gap, it's a promising architecture, not a validated one.

---

More posts: https://caio.theodoro.dev/blog.md · About the author: https://caio.theodoro.dev/about.md · Machine-readable index: https://caio.theodoro.dev/llms.txt
