Where the Money Goes Per Reply
A reply costs a fraction of a cent to a few cents, depending on model size and length. This small per-message cost adds up across millions of users, explaining why free services target heavy users and how costs shape caps, queues, and model selection.
Most debates about free AI chats are actually about a number rarely stated. That number isn't secret, and it explains the design of nearly every product in this category.
On this page
The Cost Is Hardware Time
Generating text occupies a graphics processor for the duration of the reply. Since this hardware is bought or rented by the hour, the cost of a reply is simply the fraction of an hour it consumes. Longer replies and larger models cost proportionally more, which is why both are the first limitations applied to free tiers.
Reading Is Cheaper Than Writing
Processing your question and the conversation history costs less per word than generating the response. This is why services are more generous with long inputs than with long outputs, and why a lengthy conversation increases costs gradually rather than suddenly.
Fixed Costs Are Separate
Idle hardware still incurs costs, and demand fluctuates throughout the day. A service must either buy enough capacity for peak times or accept queuing during those peaks. This choice is a financial decision, visible to you as a waiting room.
Why the Number Shapes the Product
At a fraction of a cent per reply, a light user costs almost nothing, while a heavy user costs real money. Consequently, every free service is designed around the heavy user: the usage cap, the smaller model, and the slower queue. Understanding this makes the design logical rather than arbitrary.
The questions free things deserve
Do costs decrease over time?
Yes, sharply, and demand rises to meet them. The cost per reply has dropped significantly, yet total spending has not decreased because models and conversations have grown in size and complexity.
Is a smaller model much cheaper?
Often by a factor of 10 or more. This is why silently switching to a smaller model is such a common method for controlling costs.
Does my conversation length matter?
Yes. Earlier messages are resent with each new request, meaning a long conversation costs more per reply than a short one.
Start a conversation and observe how the service behaves as usage grows.
Open the free chat