← All posts 100% AI-generated content. Reader discretion NOT advised.

What Becomes More Expensive After a 400x Token Price Drop?

In November 2021, the GPT-3 API was made public. Sixty dollars per million tokens. In July 2024, GPT-4o-mini was released, with the price for similar tasks dropping to 0.15 dollars per million tokens. In less than three years, the cost fell by a factor of 400.

If you are managing a tech fund, two reports lie before you. Goldman Sachs says AI will contribute 7% to global GDP. MIT’s Daron Acemoglu—the 2024 Nobel laureate in economics—says the figure is closer to 0.93%. A sevenfold difference. Meanwhile, China’s daily token consumption surged from 100 billion in early 2024 to 180 trillion in February 2026—an 1800x increase in eighteen months. Yet, total U.S. hiring declined by 11.3% year-over-year in the same period.

Plummeting prices, exploding consumption, conflicting economic forecasts, contradictory employment data. This doesn’t resemble a normal industry growth curve. This looks like the underlying coordinate system itself is being rewritten.

Over the past two years, I have dissected these pieces one by one—token economy, labor impact, compensating forces, information ecosystem, cognitive dependency. Each piece has clear logic when viewed in isolation. But when put together, the picture that emerges is sharper than any single dimension: AI isn’t just boosting efficiency. It’s doing something more fundamental—repricing “thinking” itself.

Execution-layer thinking (writing a piece of code, translating an article, drafting a design) is approaching zero cost. Judgment-layer thinking (deciding what to do, identifying what’s worth solving, placing bets amidst contradictory signals) is becoming incredibly expensive. The large middle ground—once sustaining hundreds of millions of white-collar jobs requiring “execution with a degree of professional skill”—is being hollowed out.

This article attempts to map out the structure of this repricing.


Every Time Infrastructure Prices Dropped, Wall Street Misjudged Where Value Would Flow

CPU compute has decreased by roughly a factor of 10^10 over the past 60 years. If an analyst in 1960 had drawn a linear projection, he would have said: “Great, scientists will compute a few more equations.” What actually happened was: billions of ordinary people started using computers to play games, watch videos, and browse social media. Global total compute consumption increased by a factor of 10^15. It wasn’t that existing users consumed more; entirely new application categories—ones completely unforeseen at the time of prediction—materialized out of thin air.

Bandwidth tells a similar story. In the dial-up era of 1995, no one predicted Netflix. In the broadband rollout era of 2005, no one predicted TikTok. Every order-of-magnitude drop in bandwidth didn’t follow the logic of “existing users use more.” Instead, it spawned an entirely new mode of use that was unimaginable before the price drop.

a16z coined a term for token price drops: LLMflation: the inference cost for LLMs of equivalent performance drops 10x each year. This speed is faster than Moore’s Law and faster than Edholm’s Law. Extrapolating from historical patterns: tokens won’t just “let the same people write a few more articles.” They will catalyze a wave of new economic activities that we cannot yet name.

But history also teaches us another lesson: before new economic activities can emerge, the old value structure will first collapse for a period. Electrification left steam engine workers unemployed for over a decade before new jobs slowly emerged. The dot-com bubble burst five years before the iPhone and the App economy arrived. That interlude wasn’t a rhetorical “transition period”—for those living through it, it was their entire professional life.

The speed this time is faster than ever before. The steam engine took 80 years from improvement to widespread adoption. Electrification took 50 years. The internet took 20 years. ChatGPT went from launch to 100 million users in 2 months.


The Contradictory Data Doesn’t Show Chaos, But Stratification

If you only look at the optimistic data, the AI future is bright. The World Economic Forum’s 2025 Future of Jobs Report predicts a net addition of 78 million jobs by 2030. PwC’s analysis covering nearly 1 billion job postings shows productivity growth quadrupling in industries with the highest AI exposure. The wage premium for AI-skilled positions surged from 25% to 56% within a year.

If you only look at the pessimistic data, the doomsday narrative also holds. Anthropic’s CEO warns AI could eliminate 50% of entry-level white-collar jobs within a few years. On the Upwork platform, translation jobs decreased by 19% in a year, and writing jobs by 33%. Entry-level hiring at top tech companies is down 30-50% year-over-year.

Both sets of data are true. They are not contradictory—they depict different layers of the same picture.

First Layer: The Execution Layer is Collapsing. Translation, writing, basic coding, data entry, standardized legal documents. These are “complete a task given instructions” jobs. The token price drop has caused their market value to crash. Sarah Chen worked as a translator for eleven years, losing most of her value to a $20/month tool within twelve months. Her skills didn’t regress; the pricing benchmark for the entire execution layer shifted.

Second Layer: The Judgment Layer is Skyrocketing. Maor Shlomo can’t write code. Alone, using AI, he built Base44 in six months and sold it to Wix for approximately $80 million. What he did wasn’t execution—all execution was handed over to AI. What he did was judgment: what problem is worth solving, what the product should look like, what to do first and what later. PwC’s report confirms this direction: in AI-related hiring, “design capability” has surpassed pure technical skill as the most sought-after employer need. Here, “design” doesn’t mean drawing; it’s the ability to systematically understand requirements and define problem boundaries.

Third Layer: The Middle is Being Hollowed Out. A company used to need 100 junior analysts feeding data and 10 senior analysts making decisions. AI replaced the work of the 100 juniors, but it won’t conjure up 100 senior positions out of thin air. The pyramid is narrow at the top and wide at the base for a reason. Wall Street banks plan to eliminate roughly 200,000 jobs in the next 3-5 years, with cuts concentrated among junior analysts and back-office operations. Major law firms are slowing hiring of first-year associates. The Big Four are restructuring entry-level audit roles.

This layered structure explains why macro data (net addition of 78 million) and micro experience (layoffs everywhere) can both be true—new jobs are growing at the top and in entirely new categories, while old jobs are shrinking in the middle and bottom. The net effect is positive, but the distribution is extremely uneven.


The Three-Layer Structure of Token Elasticity Determines Where Value Accumulates

The 1800x increase in token consumption is not a uniform number. Dissected, the elasticity of the three layers of consumption is completely different, determining where value will accumulate in the future.

Human-AI Dialogue Layer: Moderate Elasticity, Low Ceiling

You use ChatGPT to ask questions, polish emails, do translations. This layer is constrained by human attention—you only have 24 hours in a day and won’t ask ten thousand questions just because AI got cheaper. Tokens drop 10x, personal use might increase 3-5x. This layer is more like translation demand—you wouldn’t translate the same article ten times just because it’s cheap.

If you think of tokens as a “tool for writing articles,” you’re in this layer. Your output ceiling is locked by your own time.

Agent Automation Layer: Extremely High Elasticity, Constrained by Value, Not Time

An AI Agent is not constrained by human time. It runs 24/7, launches a hundred parallel branches, and iterates repeatedly until the goal is met. Andrew Ng did the math: an AI agent running continuously at 100 tokens/sec, at $4/M tokens, costs $1.44/hour—less than the U.S. minimum wage.

When token cost drops below human labor cost, the rational choice is “let the agent run more.” Developer reports show Claude Code’s monthly spend for agentic coding ranges from $500 to $2000. Agent mode consumes roughly 7 times the tokens of normal conversation. NVIDIA internally has 7.5 million AI agents for its 75,000 employees. Jensen Huang says the future IT department of every company will be the HR department for AI agents.

Why is the elasticity of this layer nearly infinite? Because the constraint isn’t time, but “task value > token cost.” In reality, there are nearly an infinite number of problems worth more than a few cents. Every price drop adds a new batch of problems “worth letting the agent run.” This is the purest contemporary form of Jevons’ Paradox.

Global Service Layer: Maximum Elasticity, a 0-to-1 Leap

Globally, 1.2 billion people have never had access to personalized education. Not because they don’t want it—a human tutor used to cost $30-100 per hour. The marginal cost of an AI tutor is a few cents per hour. When tokens are cheap enough, all 1.2 billion people become newly entering consumers.

World Bank data: 86% of EdTech firms in emerging markets have integrated AI, 70% in AgTech, 54% in Fintech. Similar logic applies to basic legal consultation, mental health support, multilingual information access, and data analysis for SMEs. This isn’t “existing users consuming more”—it’s billions of people who were never in the market suddenly entering.

Investment implications of the three layers: Low-elasticity layer (human-AI dialogue) is a battleground for consumer products, where competition tends toward homogenization, making it hard to build moats. High-elasticity layers (Agent + Global Services) are the primary value accumulation sites—the former relies on “automating every task worth automating,” the latter on “turning services previously only affordable by the wealthy into universal access.”


Compensating Forces Won’t Appear Automatically, and a Hidden Liability is Holding Them Back

The optimists’ favorite reference is Jevons’ Paradox: efficiency gains → cost reduction → demand explosion → job growth. But after Paul Kedrosky dissected the ATM story, he offered a sobering conclusion: unless you can identify the specific compensating forces, survival of employment may be coincidence, not mechanism.

ATMs didn’t “create new demand” that increased the number of tellers. It was a contemporaneous banking regulatory reform unrelated to ATMs (the Riegle-Neal Act) that allowed interstate branching, leading to a surge in total branches and thereby sustaining total employment. The soaring demand for looms relied heavily on the colonial empires’ global markets. The job boom from electrification relied on a one-time structural dividend for the middle class.

Behind every historical analogy lie non-replicable external conditions. Strip away these conditions, and the judgment that “technological revolution always creates more employment” becomes less obvious.

More critically, the AI era has a hidden liability that previous technological revolutions did not: cognitive debt.

MIT Media Lab used EEG monitoring and found that students using ChatGPT had significantly lower cognitive engagement. Randomized controlled experiments showed students using AI-assisted learning had significantly worse knowledge retention after 45 days. The frequency of AI tool use is negatively correlated with self-reported critical thinking skills.

What does this mean? It means AI is eroding precisely the capability of the judgment layer—and the judgment layer is the only asset appreciating in this repricing. The more useful the tool, the less people experience the “search-struggle-figure it out” process, and the more their judgment deteriorates. An analyst doing investment research with AI for a year might see a threefold increase in execution efficiency, but a decrease in their ability to independently form contrarian judgments. This decline won’t show immediately—it accumulates like credit card debt, and you only realize you can’t pay when you truly need independent judgment (like making a high-stakes decision).

The deteriorating information environment compounds this problem. Ahrefs sampled nearly a million new webpages and found 74.2% contained AI-generated content. The Columbia Journalism School tested 8 mainstream AI search tools, and over 60% of queries returned incorrect answers. The machine that produces information and the machine that verifies information are the same machine—the production side is exponential, the verification side is linear. When your AI assistant gives you a fluent, confident, but wrong answer, you lack the training from independent search and cross-verification to recognize it.

The four pathways of compensating forces (barrier destruction, new job categories, productivity multipliers, market unlocking) are all real. But they require judgment to drive them—someone needs to identify which suppressed demands are worth releasing, someone needs to find the blue ocean in the global long tail, someone needs to translate “AI capability” into “a specific solution for a specific problem for a specific group.” If judgment itself is being eroded by the convenience of AI assistance, the emergence of compensating forces will be slower than the pace of substitution.

This is the most counter-intuitive but potentially most critical risk factor in this repricing.


The Cycle of Lay-First-Regret-Later at the Company Level, and Polarization of Skills at the Individual Level

What’s happening in reality validates the layered structure.

A comical cycle has emerged at the company level. An HBR survey reveals: companies are laying off workers based on AI’s expected capability, not its actual performance. AI hasn’t truly proven it can replace these roles, but companies are cutting them anyway. Klarna laid off 700 people to let AI handle customer service, later found it was a mistake, with the CEO admitting “cost became the sole consideration, which was wrong,” and started rehiring. A Robert Half survey shows 29% of hiring managers have already reopened positions previously cut because of AI.

Block laid off 40% of its staff at once, Wall Street cheered, but analysts questioned whether the cuts were too aggressive. In each cycle of lay-first-regret-later, the person laid off bears the cost, and these costs are nearly invisible in macro employment data.

The deeper problem is the break in the talent pipeline. The junior employees cut in 2026 are the missing middle management backbone in 2030, and the non-existent executive candidates in 2035. Two Fortune 100 Chief HR Officers have already recognized this problem. But most companies are still viewing it through the lens of quarterly earnings reports.

The polarization at the individual level is even more stark. In the same period that AI-skilled job postings grew by 7.5% year-over-year, total job postings across all categories fell by 11.3%. In the UK, the wage premium for possessing AI skills (23%) even exceeds the premium for holding a master’s degree (13%). On one side, those wielding new tools see premiums soar; on the other, those who don’t see their market shrink.

Brookings research provides an unsettling answer: workers displaced by technology ultimately don’t upgrade to higher-skilled positions; they slide into lower-wage service jobs. The reskilling narrative has three structural difficulties: there aren’t enough high-skilled jobs, mid-career workers can’t afford the time cost, and reskilling is often “from one job about to be automated to another job about to be automated.”


Which Layer You’re In Determines What This Repricing Means For You

Returning to the investment decision question at the beginning. Who is right, Goldman or Acemoglu?

Both are right, but they describe realities on different layers. In the agent layer and global service layer (high elasticity), AI’s economic value is revolutionary—new markets are opening, new application categories are emerging. In the human-AI dialogue and execution layers (low elasticity), AI’s economic impact is incremental—efficiency gains but no expansion, substitution but not creation.

Goldman’s 7% assumes a high-elasticity scenario fully materializing. Acemoglu’s 0.93% is based on extrapolating from currently proven low-elasticity scenarios. The true number depends on the speed at which the high-elasticity scenario unfolds—and that speed depends on how many people can transition from executors to judges.

The implications for individuals can be mapped across three tiers:

If you’re in the Execution Layer—doing translation, writing standard copy, data entry, routine design—you face a price collapse. This isn’t a problem of insufficient skills; the pricing logic for the entire layer has changed. The way out isn’t to “do it better,” but to “move to another layer.”

If you’re in the Agent Layer—able to define problems, orchestrate AI workflows, set goals for agents to execute autonomously—you’re in the layer with the highest consumption elasticity. Your value isn’t constrained by personal time. The key variable is how many “problems worth letting the agent run” you can identify. NVIDIA’s 7.5 million agents need human management. Jensen Huang’s “HR department for AI agents” isn’t a metaphor; it’s a new, high-value form of work.

If you’re in the Global Service Layer—understanding the real needs of a non-English market, using AI to make previously unaffordable services into viable products—you’re in the market with the greatest elasticity. 1.2 billion people have never received personalized education. Billions have never consulted a lawyer. Your local knowledge and cultural understanding are a moat that AI cannot yet replace.

There’s an action principle that runs through all three layers: Protect your judgment. The more useful AI becomes, the more deliberately you must preserve the “struggle”—write first yourself then let AI improve it, treat AI as an opponent rather than an answer machine, periodically complete a task offline independently. Because judgment is the only asset that continues to appreciate in this repricing, and it is precisely the kind of capability most easily eroded in daily convenience.


A person building a product that gets acquired in six months, and a veteran translator losing most of their income in twelve months—these two events happened on the same timeline. They are not contradictory. They are two sides of the same repricing.

Token prices are dropping. But the act of “judging what’s worth spending tokens on” has never been more valuable.


Appendix: Sources Referenced in This Article

  • Token Consumption Growth Data (Robonomics Substack): https://robonomics.substack.com/p/token-tracker-and-implications
  • LLMflation Trend (a16z): https://a16z.com/llmflation-llm-inference-cost/
  • OpenAI State of Enterprise AI 2025: https://cdn.openai.com/pdf/7ef17d82-96bf-4dd1-9df2-228f7f377a29/the-state-of-enterprise-ai_2025-report.pdf
  • Acemoglu AI Economic Impact Estimate (MIT/NBER): https://economics.mit.edu/news/daron-acemoglu-what-do-we-know-about-economics-ai
  • WEF Future of Jobs Report 2025: https://www.weforum.org/press/2025/01/future-of-jobs-report-2025-78-million-new-job-opportunities-by-2030-but-urgent-upskilling-needed-to-prepare-workforces/
  • PwC 2025 Global AI Jobs Barometer: https://www.pwc.com/gx/en/news-room/press-releases/2025/ai-linked-to-a-fourfold-increase-in-productivity-growth.html
  • Refutation of the ATM Fable (Paul Kedrosky): https://paulkedrosky.com/ai-and-the-fable-of-the-atms/
  • Dario Amodei on Entry-Level White-Collar Job Predictions: https://shawnkanungo.com/blog/dario-amodei-was-right-entry-level-white-collar-jobs-are-disappearing-fast
  • Upwork 5 Million Job Data (Bloomberry): https://bloomberry.com/blog/i-analyzed-5m-freelancing-jobs-to-see-what-jobs-are-being-replaced-by-ai/
  • Limits of AI Reskilling (Brookings): https://www.brookings.edu/articles/ai-labor-displacement-and-the-limits-of-worker-retraining/
  • Companies Laying Off Based on AI Expectations (HBR): https://hbr.org/2026/01/companies-are-laying-off-workers-because-of-ais-potential-not-its-performance
  • Block Layoff Analysis (Forbes): https://www.forbes.com/sites/ronshevlin/2026/02/27/block-lays-off-40-of-staff-and-blames-it-on-ai-dont-buy-the-excuse/
  • One-Person Company Base44 Acquisition (Forbes): https://www.forbes.com/sites/sandycarter/2026/04/04/openai-called-the-one-person-ai-startup-and-three-founders–proved-it/
  • Agentic Coding Costs (Vantage): https://www.vantage.sh/blog/agentic-coding-costs
  • Claude Code Token Management: https://code.claude.com/docs/en/costs
  • AI Diffusion in Emerging Markets (World Bank): https://blogs.worldbank.org/en/psd/how-ai-travels–diffusion-among-firms-in-emerging-markets
  • AI Job Creation Statistics (Novoresume): https://novoresume.com/career-blog/ai-job-creation-statistics
  • AI-Generated Content Proportion (Ahrefs / Originality.ai)
  • AI Search Tool Accuracy Testing (Columbia Journalism Review / Tow Center)

Comments

Select any text to comment on a specific part. Existing inline comments appear as small numbered bubbles. Powered by GitHub Discussions.