A burst of corporate interest in “tokenmaxxing”—pushing generative AI systems to use as many tokens as possible—appears to be fading as companies see higher bills without proportional productivity improvements, according to reporting that draws on multiple executives and analysts. The trend gained attention in early months, with some tech leaders treating high token usage as a sign of strong performance and rolling out internal contests or supportive remarks about maximizing token consumption. Over time, however, enterprises reassess return on investment as token costs escalate, with estimates cited that show spending doubling rapidly for large developer workforces.
Consultants and executives describe shifting from defaulting to the most capable, most expensive models for all tasks toward more disciplined use and “model routing,” where simpler requests are directed to cheaper models and complex work uses stronger ones. Some critics also raise concerns that customers effectively pay twice—through token spending and by sharing proprietary data. Meanwhile, interest in lower-cost alternatives, including open-source models from companies such as Moonshot and Zhipu, grows as teams try to control subscription and usage costs.
Sources characterize tokenmaxxing as a short-lived corporate fad rather than a lasting best practice.