It’s hard to tell how many tokens you’re getting, and it changes on you for mysterious reasons:
Before the change, somewhere between 30% and 67% of requests included a thinking block, and output averaged 250 to 960 tokens per request. After it, 95% to 100% of requests included thinking, and output rose to 2,500 to 3,000 tokens per request in the first hours. A clear example of the API changing without any user involvement.
And, then, this is on-top of teams just having problems estimating token usage. Below, an AI summary of IDC in that:
🤖: The cost failure in agentic AI is a governance failure, not a modeling one: teams underestimate input tokens by three to five times because they model the user message and not the full prompt, inference-only cost models undercount true total cost of ownership by 30% to 60%, and by the time finance sees the invoice three to four weeks have passed.
Leave a Reply