AI Subscription Jargon, Translated Into Plain English (2026)
Rolling windows, tokens, compute, weekly caps, reasoning modes — what the words on AI pricing pages actually mean, in plain English, with one metaphor each.
So — which one should you buy?
TL;DR. AI pricing pages are written in a dialect. Here's the dictionary: a rolling window is a bucket that slowly refills (no midnight reset); a token is the AI's word-piece unit (~1,000 tokens ≈ 750 words); compute is the actual work behind a message, which is what you're really buying; a weekly cap stops round-the-clock heavy use on top of the hourly window; and reasoning modes burn allowance much faster because the AI is doing more work per message. Once you know these five, every AI pricing page reads the same.
"I still don't understand the AI jargon." A reader said exactly that to our site guide while standing on a page full of numbers — and they were right to. The words on AI pricing pages ("rolling window," "compute-based limits," "token allowances") aren't hard because the ideas are hard. They're hard because nobody translates them. Here's the translation, one term at a time, each with a single metaphor you can keep.
"Rolling window" — a bucket with a slow drip
Your message allowance doesn't reset at midnight. It refills continuously on a timer — usually five hours. Picture a bucket with a steady drip refilling it: every message you send takes a scoop out; wait a while and the level climbs back on its own.
Two consequences people miss: there is no magic reset moment you can wait for (the refill is gradual, tied to when you used the messages), and unused allowance never banks — a quiet morning doesn't earn you a bigger afternoon. Both Claude and ChatGPT work this way in 2026.
"Token" — the kilowatt-hour of AI
A token is the AI's unit of text — a word, or a piece of one. Rule of thumb: 1,000 tokens ≈ 750 words. You never buy tokens directly on a consumer plan, but they're the meter running behind everything: what you type, what you upload, and what the AI writes back all get counted in tokens, the way your electric company counts kilowatt-hours. When a vendor says limits are "token-based," it means the meter measures the amount of text processed, not the number of times you hit send. (Deeper dive: tokens and context windows, explained.)
"Compute" — what you're actually buying
Here's the concept that makes everything else click: vendors don't sell messages, they sell work. A two-line question makes the AI do a sliver of work. Uploading a 200-page PDF and asking for analysis makes it do orders of magnitude more. Both count as "one message."
That's why no honest pricing page promises a fixed message count anymore — two messages can differ in cost by 100×. When you see "limits vary with usage," it isn't evasion; it's the vendor saying we meter the work, not the sends. The practical move: short chats and small attachments stretch your plan; giant uploads and marathon conversations drain it.
"Weekly cap" — the second ceiling
On top of the rolling window, most plans add a weekly ceiling. Why both? The rolling window stops short bursts; the weekly cap stops someone from sitting exactly under the hourly limit around the clock, all week. Most people never touch it. If you've ever seen "you've reached your weekly limit" while the hourly meter looked fine — that's this, and it means your overall volume, not your pace, hit the line.
"Reasoning mode" / "Deep research" — the gas-guzzler setting
These modes make the AI think longer and search deeper before answering — and they consume allowance much faster than normal chat, because each message triggers far more work (see "compute," above). They're worth it for genuinely hard tasks. The mistake is leaving them on for routine questions, which is the single most common way people burn a day's allowance by lunch.
"Model tiers" — light, medium, and heavy machinery
Every vendor now offers a lineup of models — lighter ones that are fast and cheap to run, heavier ones that are slower, smarter, and hungrier. Your allowance stretches or shrinks depending on which one you pick: the same plan might give you hundreds of light-model messages or a few dozen heavy-model ones. (OpenAI even publishes its message estimates as ranges per model now.) The strategy is the same everywhere: default to the lighter model, save the heavy one for work that needs it. If you're not sure which Claude model fits a task, this picker answers it in two questions.
The limits of the metaphors
One honest caveat: all of these systems flex. Vendors adjust limits with demand, publish ranges instead of numbers, and change the rules more often than the pricing pages get rewritten. Treat any specific figure as a dated snapshot — including ours — and treat the concepts above as the stable part. The concepts haven't changed in two years; the numbers change monthly.
That's the whole dialect. Rolling window: refilling bucket. Token: kilowatt-hour. Compute: the work, which is what you're buying. Weekly cap: the second ceiling. Reasoning modes: gas-guzzlers. Model tiers: light vs heavy machinery. Now go re-read the actual numbers — or the ChatGPT and Claude plan breakdowns — and notice they've stopped being confusing.
More terms, one paragraph each, live in the AI glossary.
So — which one should you buy?
Set up AI for your job — free, in about 2 minutes
Pick your profession and get your first working AI tool, a step-by-step guide, and a $0 plugin to take home. No credit card.
Get my free setupFrequently asked questions
What is a rolling window in AI subscriptions?+
A rolling window means your message allowance refills continuously on a timer — usually five hours — instead of resetting at midnight. Think of a bucket with a slow steady drip refilling it: use messages and the level drops, wait and it climbs back. There's no magic reset moment, and unused allowance never banks up.
What does 'compute' mean on an AI pricing page?+
Compute is the actual work the AI does to answer you, and it's what vendors really meter. A two-line question costs a sliver; uploading a 200-page PDF or turning on a reasoning mode costs orders of magnitude more. That's why vendors stopped promising fixed message counts — two 'messages' can differ in cost by 100x.
What is a token in plain English?+
A token is the AI's unit of text — a word or piece of a word. Roughly 1,000 tokens is about 750 words. Vendors meter usage in tokens the way a utility meters electricity in kilowatt-hours: it's the billing unit behind the scenes, and everything you send and receive gets counted in it.
Why is there a weekly cap on top of the 5-hour limit?+
The rolling window stops short bursts; the weekly cap stops sustained heavy use. Without it, someone could sit exactly under the 5-hour limit around the clock all week. Most people never touch the weekly ceiling — it exists for the heaviest users, and it's why you occasionally see 'you've reached your weekly limit' even when the hourly window looks fine.
Related Guides
AI Data Governance: Which Data Goes Into Which Tool
A practical AI data governance model for data strategy and governance leads: a data-class-by-tool-tier matrix, a usable AI use register, the vendor retention questions worth asking, and where governance and AI spend data overlap.
Build a Personal AI Operating System in Claude, Step by Step
The complete setup: exact prompts and files to build a working personal AI OS in Claude Cowork in under an hour — profile wizard, card-menu command center, skills, always-on guard, and one scheduled routine.
How to Run a Solo AI Agency on One Operating System
Acquisition, scoping, delivery, retainer reporting, and content are five jobs held by one person. Here's how to run them as a single AI operating system instead of five disconnected tools.