A production trace with 6.12 billion rows just became public, and systems researchers are paying attention. On October 8, 2026, coverage continued to circulate around a dataset release first announced in late September: Chutes AI, a decentralized inference network, teamed up with a Harvard researcher and collaborators at the University of Chicago to publish a year of real LLM serving metadata. No prompts. No model outputs. Just the plumbing of how 9,174 models actually got served to 314,970 users, recorded request by request.

The dataset, titled “A Year in LLM Serving,” landed as a single 91 GB Parquet file hosted on a Harvard-controlled bucket, with a companion paper on arXiv and a public landing page at data.agentic-system.org. For an industry that usually guards its traffic numbers like trade secrets, the size and openness of the release stands out. It is also a useful stress test for a question that keeps resurfacing in AI infrastructure circles: who actually gets to study how people use language models at scale, and on what terms.

What Chutes and Harvard Actually Released

Strip away the framing and the release is simple to describe: a year-long request log from a working inference platform, stripped of content and handed to researchers. According to the project’s own announcement and reporting from Tao Media, the trace covers 6,122,413,756 individual LLM inference requests spanning 9,174 distinct models. The observation window, per a technical write-up from ibl.ai, runs from April 11, 2025, through April 12, 2026, almost exactly 365 days of production traffic.

Scale alone sets this apart from most public LLM research data. The trace logs 35.8 trillion input tokens processed across 875,921 serving instances, numbers that come directly from Chutes’ own release notes. Most academic LLM datasets top out in the millions of interactions. This one runs three orders of magnitude higher, and it does so by logging infrastructure events rather than conversations.

That distinction matters. The dataset explicitly excludes prompts, model responses, and tool calls, according to ibl.ai’s review of the release. What remains is closer to a server access log than a chat transcript: timestamps, token counts, cache hits, and model identifiers, with user identity reduced to rotating anonymized IDs. Researchers get to see the shape of demand without seeing what anyone actually asked the models.

Inside the 91GB Trace: Schema and Scale

Reporting from ibl.ai lists the file’s structure at twelve metadata fields per request. The documented fields cover an invocation ID, the function being called, the chute (serving unit) ID, a user ID, a rehash round marker, an instance ID, start and completion timestamps, and token-level detail including input, output, and cached token counts, plus a field recording time to first token.

FieldWhat it recordsWhy it matters for research
invocation_idUnique identifier for a single requestLets researchers de-duplicate and join records
function_nameThe type of call being servedEnables grouping requests by workload category
chute_idThe serving unit (model deployment) handling the requestSupports per-model demand analysis across 9,174 models
user_idAnonymized, rotating user identifierAllows cohort analysis without exposing identity
rehash_roundMarker tied to Chutes’ internal routing logicUseful for studying load-balancing behavior
instance_idIdentifier for one of 875,921 serving instancesSupports infrastructure capacity-planning studies
started_at / completed_atRequest start and finish timestampsDrives arrival-rate and latency modeling
input / output / cached token countsToken volume per request, including cache reuseCore input for prefix-cache and cost-modeling research
time to first tokenLatency before the first output token streamsBenchmark for serving-system responsiveness

The combination of cache-hit data and time-to-first-token readings is what caught the eye of systems researchers first. Prefix caching, where a serving system reuses computation from earlier requests sharing a common prompt prefix, is one of the biggest open problems in production LLM serving. A year of real cache-hit behavior across thousands of models gives researchers something they have rarely had: ground truth at scale, rather than simulated workloads built from guesswork.

Chutes AI: The Network Behind the Data

Chutes is not a household name the way OpenAI or Anthropic are, and that is part of the story. It operates as an open inference network built on Bittensor Subnet 64, a decentralized compute marketplace rather than a conventional centralized cloud API. Instead of one company running all the servers, Chutes coordinates traffic across a distributed set of operators, described in the dataset materials as “chutes,” who get paid in TAO, Bittensor’s native token, for serving inference requests.

Tao Media’s coverage notes that Chutes recently named Christian Hyland as CEO as the platform expanded its AI compute footprint. Separate commentary circulating on X around the same period pegged Chutes’ throughput at roughly 27 billion tokens served per day, though that figure comes from third-party analyst commentary rather than an official Chutes disclosure and should be read as an informal estimate rather than a confirmed metric.

What makes Chutes a useful subject for this kind of release is precisely that it is not a single frontier lab. A decentralized network spanning many operators and 9,174 models captures a messier, more representative slice of real-world inference demand than a single company’s flagship chatbot traffic would. Long-tail models, not just the handful of famous ones, show up in the data.

The Research Team Behind “A Year in LLM Serving”

The academic side of the project centers on Juncheng Yang, an assistant professor at Harvard University, working alongside co-authors including William Nixon, Jon Durbin, Florian Standhartinger, and Haryadi S. Gunawi, according to ibl.ai’s account of the paper’s authorship. Gunawi’s involvement brings a University of Chicago systems research connection into the project, positioning the work at the intersection of academic systems research and commercial inference operations.

Yang announced the release directly. “Announcing one year of LLM inference metadata traces, with 6.12 billion requests,” Yang wrote on X, in a post that also framed the project’s goal: “We hope this dataset can support research on real-world LLM serving workload understanding, system design and infrastructure optimization,” he said, according to the original announcement thread.

The OpenTensor Foundation, which stewards the Bittensor network that Chutes runs on, amplified the release to its own community. “A Harvard research team and @chutes_ai just released a public dataset covering one year of real-world LLM inference on Chutes: 6.12B requests across 9,174 models,” the organization posted on X, adding that “technical usage data from a Bittensor subnet is now open to the wider AI research community.”

What’s Missing: Privacy and Open Questions

The release describes its anonymization approach in broad strokes: rotating user IDs, no prompt or response content, no tool-call logs. What it does not spell out, at least in the materials reviewed, is the underlying mechanism. There is no published detail on whether user IDs are hashed, salted, or pseudonymized through some other method, how often rotation occurs, or whether rare users and rare models were suppressed to reduce re-identification risk.

There is also no public mention of an institutional review board sign-off or a formal privacy audit tied to the release, based on the reporting available. That gap does not necessarily mean one did not happen. It does mean outside researchers cannot currently verify the strength of the anonymization guarantee from public documentation alone, a point ibl.ai’s analysis raised directly when it cautioned against framing the project simply as a “Harvard release” rather than a joint academic-commercial effort built on a commercial inference provider’s production logs.

Metadata-only release strategies are generally viewed as lower-risk than raw-content releases, since removing prompts and completions eliminates the most obvious privacy exposure. But metadata is not risk-free. Request timing, token volume patterns, and model-selection habits can, in combination, still narrow down who a user might be, particularly for low-traffic models with small user cohorts. Whether that risk was modeled formally is one of the open questions this release leaves on the table.

How It Compares to Other Public LLM Usage Datasets

Public datasets describing how people actually use language models are rarer than benchmark leaderboards, and the handful that exist tend to specialize in different layers of the stack. The Chutes trace is the first at this scale to focus purely on serving infrastructure rather than conversation content.

DatasetWhat it containsScaleBest suited for
Chutes “A Year in LLM Serving”Request-level serving metadata: timing, tokens, cache hits, instances6.12B requests, 9,174 models, ~1 yearSystems research: caching, scheduling, capacity planning
LMSYS Chatbot Arena dataUser-submitted prompts and pairwise model preference votesVaries by releaseModel evaluation and head-to-head comparison research
WildChatReal user conversations with chat modelsVaries by versionPrompt behavior, safety, and language-use research
Anthropic Economic IndexAggregated task and occupation-level usage patternsPublished aggregate reportsEconomic and labor-market impact analysis

The split is clean once you line the datasets up side by side. LMSYS Chatbot Arena and WildChat expose content, which makes them valuable for studying what people ask models and how models respond, but both work at a smaller scale and carry heavier privacy considerations because the text itself is sensitive. The Anthropic Economic Index sits at a different altitude entirely, aggregating usage into occupational and task categories rather than publishing request-level records at all.

Chutes’ trace fills a gap none of the three cover well: how the serving layer itself behaves under real, sustained, multi-model load. That is a narrower research question than “what do people ask chatbots,” but it is one that infrastructure teams at every major AI lab have been trying to answer internally for years, usually with data they never publish.

Why This Matters for AI Infrastructure Teams

Every inference provider runs on assumptions about how traffic behaves: how bursty request arrivals are, how much benefit prefix caching delivers in practice, how demand splits between a handful of popular models and a long tail of niche ones. Those assumptions normally come from internal telemetry that never leaves the building. Public traces are usually synthetic, built to approximate real traffic rather than measure it.

A real trace spanning 9,174 models changes the baseline. Researchers building new scheduling algorithms, cache-eviction policies, or autoscaling systems can now validate against a dataset drawn from actual production behavior rather than a workload generator calibrated on guesses. That has downstream value for GPU capacity planning too: cloud providers and inference platforms sizing fleets of accelerators depend on realistic demand curves, and most of the public ones in circulation are years out of date relative to how agentic workloads behave today.

The timing lines up with a broader industry conversation about open infrastructure. Decentralized or open inference networks like Chutes have spent the last two years arguing that distributed compute marketplaces can match centralized clouds on cost and reliability. Publishing a year of unsampled production data is also a credibility move. It is hard to dismiss a network’s claims about scale when the claims come with 6.12 billion rows of receipts attached.

A Short History of Public LLM Usage Data

Public visibility into how language models actually get used has always lagged behind the pace of model releases themselves. Early chatbot research leaned on small, hand-collected conversation sets. LMSYS changed that by crowdsourcing pairwise comparisons at Chatbot Arena, which gave the field its first large, continuously updated signal on model preference. WildChat extended that further by publishing full conversations contributed by users of a free chat interface, offering a window into unscripted, real-world prompting behavior.

What none of those releases captured was the operations side of serving: how requests arrive, how caching performs, how infrastructure strains under real load. Companies have published papers describing serving systems, including details about scheduling and batching strategies, but almost never with the underlying trace attached. The Chutes and Harvard release breaks from that pattern by publishing the receipts rather than just the conclusions, which is why systems researchers reacted to it differently than they typically react to a new conversation dataset.

Industry Reaction

Coverage of the release split along predictable lines. Tao Media’s write-up treated it as a straightforward infrastructure research story, framing the dataset as a resource for studying LLM serving, caching, and load balancing. The OpenTensor Foundation leaned into the Bittensor angle, presenting it as proof that decentralized compute networks can produce research-grade output alongside commercial service.

ibl.ai’s analysis took a more skeptical tone, situating the release inside a wider argument about agent infrastructure and accountability. Its piece argued that even thorough metadata releases like this one still leave a structural gap: twelve fields of serving metadata say nothing about what an AI agent actually did on a user’s behalf, a distinction the piece treated as increasingly important as more inference traffic shifts from single chat turns to multi-step agentic workflows.

Competitive Landscape: Open Networks vs. Centralized Clouds

Centralized cloud providers running inference at scale, Amazon, Microsoft, Google, and the frontier labs themselves, sit on serving traces that almost certainly dwarf Chutes’ 6.12 billion requests. None of them has published anything comparable. That asymmetry is the quiet subplot of this release: a decentralized network built on a cryptocurrency-adjacent compute marketplace just did something none of the trillion-dollar cloud incumbents have chosen to do.

It is worth being precise about why that gap exists. Centralized providers face strict enterprise confidentiality obligations tied to customer contracts, and they compete directly on serving efficiency, so a detailed trace doubles as a competitive disclosure. Chutes operates in a different competitive position: openness itself is the pitch, since a decentralized network needs public trust and researcher goodwill to justify routing traffic through third-party operators rather than a single trusted vendor.

That dynamic puts pressure in an unusual direction. Rather than open-source advocates pushing labs to release weights, this release pushes on a different axis entirely: operational transparency. Expect that framing to show up in future comparisons between decentralized inference marketplaces and traditional cloud GPU rental, not just in model-quality benchmarks.

Predictions: Where This Goes Next

  • Expect at least one other decentralized or mid-size inference provider to announce a comparable trace release within the next two quarters, following Chutes’ credibility playbook.
  • Systems conferences including MLSys and NeurIPS workshops are likely to see submissions built directly on this trace, given how rare request-level serving data at this scale has been until now.
  • Privacy researchers will probably publish a follow-up critique examining the anonymization methodology in more depth, since the current public documentation does not detail the specific technique used for user ID rotation.
  • Centralized cloud providers will face more public questions about releasing comparable, even partial or sampled, serving traces, especially from academic partners seeking similar data for publication.
  • Derivative tools, such as synthetic workload generators calibrated on this trace, are likely to appear within a year, mirroring how Chatbot Arena and WildChat each spawned downstream research tools built on top of the original release.

Risks and Open Criticism

Not every reaction to the release has been positive, and the caution is worth taking seriously rather than treating this as an unqualified win for open data. The core criticism, raised most directly by ibl.ai, is that metadata about serving infrastructure, however large, does not substitute for actual visibility into what AI agents do when they call a model. As more inference traffic comes from autonomous agents chaining multiple calls together rather than a human typing a single question, serving-layer metadata captures less and less of the behavior that actually matters for safety and accountability.

There is also a framing risk in how the release gets described. Characterizing it purely as a Harvard dataset undersells the fact that the underlying traffic comes from a commercial inference network with its own incentives to demonstrate scale and attract users. Neither framing is wrong exactly, but conflating them glosses over a distinction that matters when evaluating how representative or how neutral the data really is.

Frequently Asked Questions

What is the Chutes and Harvard LLM dataset?
It is a public dataset of 6.12 billion LLM inference requests collected over roughly a year from Chutes, a decentralized inference network, and released alongside researchers from Harvard and the University of Chicago under the title “A Year in LLM Serving.”

Does the dataset include actual prompts or model responses?
No. According to reporting on the release, the dataset contains serving metadata only, including timestamps, token counts, and cache information, and explicitly excludes prompt text, model outputs, and tool-call logs.

How large is the dataset file?
The release is described as a single Parquet file of approximately 91 GB, hosted on a Harvard-affiliated storage bucket, with a companion landing page at data.agentic-system.org.

What is Chutes AI?
Chutes is an open, decentralized LLM inference network built on Bittensor Subnet 64, routing inference traffic across distributed serving operators rather than a single centralized provider.

Who led the research behind the dataset?
Juncheng Yang, an assistant professor at Harvard University, led the public announcement, with co-authors including William Nixon, Jon Durbin, Florian Standhartinger, and Haryadi S. Gunawi, who brings a University of Chicago systems research connection to the project.

How does this compare to datasets like LMSYS Chatbot Arena or WildChat?
Those datasets focus on conversation content and model preference, while the Chutes trace focuses purely on serving infrastructure: request timing, token volume, and cache behavior across 9,174 models, with no conversational content included.

Is the data fully anonymized?
The release uses rotating, anonymized user IDs and excludes content entirely, but public documentation does not detail the specific anonymization method, rotation frequency, or whether a formal privacy or institutional review was conducted.

Why does a dataset like this matter for AI infrastructure research?
It gives researchers real, large-scale evidence for how caching, scheduling, and load-balancing actually behave in production, replacing synthetic workload assumptions with a year of measured traffic across thousands of models.