Compute, Data, and the Shape of the Next Generation
By Telasian Labs
Abstract
Scaling laws met the data wall in 2024-2025. The next generation of frontier models is characterized by synthetic data composition, post-training compute reallocation, and reasoning at inference. This paper maps the shifting compute economics and argues the dominant lever for capability gain is no longer pre-training compute scaling but inference-time reasoning depth, with second-order implications for capital allocation across the field.
The Chinchilla-era assumption was that pre-training compute and pre-training data were the two binding constraints on frontier capability, and that scaling both together along a known curve would produce predictable capability gains. By the end of 2024, both halves of that assumption had run into limits. The available high-quality data did not scale to keep pace with the compute available. The capability gains from continuing to scale pre-training compute alone, holding data quality constant, were sublinear.
The next generation of frontier models is structured around three responses to that situation: synthetic data composition, post-training compute reallocation, and reasoning at inference time. This paper maps each of those responses, traces the resulting shift in compute economics, and argues that the dominant lever for capability gain in the next eighteen months is inference-time reasoning depth, with capital allocation implications that the field has not fully priced in.
The conclusion the paper builds toward is that the visible scaling-laws-driven capital allocation in the field is being deployed against a theory of capability that the actual frontier has already moved past. The labs that have updated their internal theory of capability are quietly building infrastructure around inference and post-training compute. The labs that have not are pouring more pre-training compute at a frontier where pre-training compute is no longer the binding lever, and the consequences of that misallocation will start to be visible in 2027.
The data wall
The high-quality public web does not contain enough tokens to feed the next generation of frontier models at the data-to-parameter ratios that Chinchilla and its successors predicted. The estimates vary, but the convergent view is that the high-quality public corpus is roughly an order of magnitude too small for the parameter counts the next-generation models would want. Synthetic data has filled some of that gap, with serious risks attached.
The structure of the data wall is not just about quantity. The high-quality slice of the public web is concentrated in a small number of source distributions, most of them already heavily represented in current training corpora. Adding more tokens from these sources produces diminishing returns. Reaching into lower-quality sources produces degraded capability on the dimensions that matter. The labs that try to scale past the wall by simply adding more pre-training tokens find themselves on the wrong side of the quality-quantity tradeoff, and the resulting models do not produce the capability gains the scaling curve predicted.
The data wall is also asymmetric across domains. There is enough high-quality data in some domains (general web, code, mathematics) to keep scaling for one or two more generations. There is not enough in others (specialized professional knowledge, niche technical domains, languages with smaller online presence) to support frontier-scale training without substantial synthetic augmentation. The capability profile of the next generation of frontier models will reflect this asymmetry, with some domains continuing to advance and others stagnating or regressing relative to user expectations shaped by the previous generation.
The synthetic data response
The lab generates additional training data from a previous-generation model or a deliberately-prompted version of the current model, filters it for quality, and includes it in the next generation's pre-training corpus. This works in measured contexts. It fails catastrophically in others. The failure mode is mode collapse: the synthetic corpus inherits the prior of the model that generated it, the next generation absorbs that prior, and capability across some axes silently degrades while benchmarks continue to improve.
The labs that take synthetic data seriously have built infrastructure for synthetic data composition that is one of the most under-discussed parts of the frontier stack. The decisions about which slice of capability to generate synthetically, what filter criteria to use, what proportion of the training corpus to allow synthetic content to occupy, are some of the highest-leverage decisions any lab makes. They are also some of the least transparent.
The mode-collapse failure mode is structural and predictable. Each generation absorbs the prior of the generation that produced its synthetic training data. The prior is not the same as the underlying capability distribution; it is the distribution of outputs the previous generation was good at producing under the prompting strategy that was used to generate the synthetic corpus. Capability dimensions outside that distribution attenuate generation over generation. The labs that take this seriously rotate their synthetic data sources, mix synthetic with fresh human-curated content at known ratios, and reserve specific capability dimensions for non-synthetic training only. The labs that do not take it seriously will ship models that are visibly stronger on the dimensions their previous generations were already strong on and silently weaker on the dimensions that did not get rotated into the synthetic corpus.
The synthetic-data attribution problem
When a frontier lab trains on synthetic data generated from another lab's model, the question of where the capability actually came from becomes structurally unanswerable. Lab A trains on synthetic output from Lab B's previous-generation model. Lab A's next generation shows capability gains in domains where Lab B was strong. Is that a Lab A capability or a transferred Lab B capability? The cross-pollination has been happening for at least two generations, and the resulting model behavior is shaped by a synthetic substrate whose composition is determined by a previous generation of models that were themselves partly synthetic. The attribution problem is not yet a legal or competitive issue, but it will be.
Post-training as the new battleground
Compute previously spent on pre-training scaling is increasingly being reallocated to post-training. Reinforcement learning from human feedback, constitutional methods, multi-stage instruction-tuning pipelines, tool use training, agentic task training. The compute budget for post-training at frontier labs has gone from a small fraction of pre-training compute to comparable or in some cases larger. The capability gains from this reallocation are real and not predicted by scaling laws focused on pre-training.
This shift is structural, not transient. The marginal capability return on pre-training compute is decreasing. The marginal capability return on post-training compute is currently high and probably has a longer runway than the current consensus assumes. Capital that follows the visible scaling laws is being misallocated relative to the actual frontier.
The post-training stack has also become substantially more sophisticated than the public conversation acknowledges. Frontier labs run multi-stage post-training pipelines with reward models trained on curated human preference data, constitutional methods that apply structured critiques during fine-tuning, tool-use training that exposes the model to agentic environments, and reasoning training that shapes how the model uses inference-time compute. The capability gains from each of these stages compound, and the labs that have built the full pipeline have a substantial capability advantage over labs that are still operating with a single-stage RLHF approach. The advantage is invisible in the public benchmarks because the benchmarks do not measure the dimensions where the post-training pipeline matters most.
Capital that follows the visible scaling laws is being misallocated relative to where the actual frontier is. The shift requires a different theory of capability, and the theories that drive most of the visible capital deployment are still anchored on a pre-training-centric view of scaling that the frontier has already moved past.
Inference-time reasoning
The largest architectural shift of the past year is reasoning at inference. Chain-of-thought scaled into multi-step inference budgets, allowing models to spend more compute per query in exchange for better performance. The capability gain is not subtle. On hard reasoning benchmarks, the same base model with extended inference-time reasoning outperforms the same model with standard inference by margins that are uncomfortably large for the previous generation of compute economics to make sense of.
The implication is that inference compute is now a first-class capability lever, not a runtime cost. The labs that build inference-time reasoning into their architecture and serving stack are buying capability gains that are not available through pre-training scaling at any reasonable cost.
The economic implications of this shift are still being worked out. A model that spends a thousand-fold more compute on a hard query than on an easy one needs a serving infrastructure that can dynamically allocate that compute, a billing model that can charge for it sustainably, and a user experience that can handle the latency variance. None of these are fully solved. The labs that solve them first will have a serving stack that supports the capability profile of inference-heavy models, and the labs that have not built that infrastructure will find themselves with strong models they cannot serve economically at the inference profiles those models actually require.
Inference-time reasoning as a new scaling axis
The scaling laws that have driven the field for the past five years describe how capability scales with pre-training compute and pre-training data. There is now a third scaling axis: inference-time reasoning depth. The capability-as-a-function-of-inference-compute curve is not yet well-mapped, but the early data suggests it has substantial runway. The labs that map this curve carefully will have a model of capability gain that the rest of the field is still operating without, and that asymmetry will compound over the next two generations.
Capital allocation
If the dominant lever for capability gain is inference-time reasoning, the capital that is being deployed into pre-training compute is partially misallocated. The right ratio of pre-training to inference compute investment is shifting toward inference, and the lag in capital allocation is producing a window where labs that move first into inference-heavy architecture get a real advantage that compounds.
This is the kind of structural shift that scaling-laws-focused capital allocation cannot price. The shift requires a different theory of capability, and the theories that drive most of the visible capital deployment are still anchored on a pre-training-centric view of scaling that the actual frontier has already moved past.
The mismatch shows up in announced capital commitments. The largest visible capital commitments in the field are for pre-training compute clusters that will come online in 2027 or 2028. These commitments were made on a theory of capability that was already incomplete when they were signed. The labs that hold them will get value from the compute, but the ratio of that value to the value of equivalent capital deployed into inference and post-training infrastructure will be substantially worse than the field-wide spreadsheets project. The capital reallocation that will eventually correct this is itself a multi-year process, and the labs that started reallocating earliest will be ahead at the point in 2028 when the rest of the field acknowledges the shift.
Three predictions for 2027
First: the dominant frontier models will spend more compute on a single hard query than they spend on training data per query during pre-training. Inference compute will routinely exceed pre-training compute on a per-token basis for the queries that matter.
Second: at least one major frontier lab will publicly disclose that its capability gains over the past twelve months came predominantly from post-training and inference-time reasoning, not from pre-training compute scaling. The acknowledgment will reshape the public narrative about how frontier capability is produced.
Third: a serious capital reallocation will follow, with inference infrastructure investment growing at a multiple of pre-training compute investment. The labs that built inference-first capability ahead of the shift will have moved decisively ahead of those that did not.
Conclusion
The compute-and-data picture that drove the field through 2024 is no longer the picture that describes the frontier. The data wall is real. Synthetic data fills the gap with structural risks attached. Post-training has become a comparable or larger share of the compute budget than pre-training. Inference-time reasoning is the dominant new lever for capability gain. Capital allocation across the field has not yet adjusted to any of these shifts in a comprehensive way.
The labs that have updated their internal theory of capability are building toward an inference-heavy, post-training-heavy frontier that looks substantially different from the pre-training-scaling frontier of two years ago. The labs that have not updated will continue to ship models on the old theory and find that their capital is producing diminishing returns. The window in which the shift can be made cleanly is open now and will not stay open through the end of 2027. After it closes, the labs that did not move will be operating at a capital efficiency disadvantage that is structurally hard to recover from.
Citation
Telasian Labs. (2026). Compute, Data, and the Shape of the Next Generation. https://telasian.com/papers/compute-data-and-the-shape-of-the-next-generation