The Time-Suck Era of AI Is Ending. Design Decides What Comes Next.

Four years and hundreds of billions a year in, daily life outside tech is unchanged. The shift from impressive to useful is a design problem, not a model one.

A
Arpy Dragffy · · 8 min read
Overview
  • Hyperscalers committed $660–690B in 2026 capex while only 13% of organizations report performing significantly better because of AI.
  • Rebecca Hinds measures 6.4 hours a week lost to botsitting; Molly Sands measures 2.4 billion hours a year lost to search inside the Fortune 500.
  • {"Three 2026 shipments change the calculus":"ChatGPT Voice driving Codex agents, Apple's on-device foundation models, and Grok Bot's chief-of-staff pattern."}
  • The constraint was never model capability. Turning impressive into useful is design work, and design is the discipline AI spent four years making look optional.

Here we are four years into the most expensive technology build-out in history, almost nothing about ordinary life has changed.

The five largest cloud providers have committed between $660 and $690 billion in capital expenditure for 2026 alone, close to double what they spent last year. In the same window, Glean's Work AI Index found the average worker burning 6.4 hours a week cleaning up after the tools that were supposed to give time back. AI has become more of a giant time-suck than useful technology.

As I sit here doom-scrolling through endless AI hype videos, I have to question what it will take for LLMs to break through and have impact in ways that matter: Improved health outcomes, deeper IRL relationships, and more satisfying shopping experiences. Sure, our access to information about all of these has improved, but AI isn't yet an applied solution outside of research settings.

2026 appears to be when the tide is turning. The barriers to adoption and actual value-creation — moving usage outside of the chat interface — are starting to fall. Three major developments: ChatGPT Voice arriving on the desktop with control over Codex agents and computer use, Apple's third generation of foundation models running natively on device, and xAI's Grok Bot handing every user a persistent cloud machine with a chief-of-staff agent routing work to specialists.

If we look at GenAI as more than data and instead as a design solution to real problems, then we can start taking this technology seriously. Here's the path forward.

So what did four years of building buy us?

Think about building a bridge. Four years in, we're finally getting the pilings right: models that follow instructions, tools that call other tools, context that survives a session, agents that finish a task without someone hovering over them. Those are foundations, not spans. Nobody drives across a footing.

The industry keeps reporting footing quality like it's traffic volume. Every quarter the benchmarks get more impressive and the release notes get grander, while the number of people whose week got measurably better stays flat. That same Work AI Index found 87% of digital workers using AI and 75% saying it makes them more productive, and only 13% of their organizations performing significantly better because of it.

You might be reading this thinking your output has never been higher, and you'd be right. The slop cannon is more loaded than ever, and anyone publishing into a feed can feel it. Volume is the one thing this era has reliably produced, and volume isn't impact, which is why the companies generating the most AI output aren't the ones reporting better results.

Look at how you booked your last trip

Travel is the cleanest test I can think of, and you can run it on yourself. Booking a trip is high-consideration, high-emotion, and drowning in information, which is the condition where a capable assistant should change your life. Think about the last trip you planned. The only thing AI changed was which box you typed the question into.

You still can't get a system to put you in the neighbourhood you'd love, weigh flights against what you care about instead of price alone, or make sure you don't miss the one experience that would have defined the trip. Those are the outcomes people want, and not one of them is a model problem. Each requires knowing a person well enough to act on their behalf, which is research and design work the industry has mostly declined to fund.

The money is real. The payoff still isn't.

None of this is speculative spending. Hyperscaler capital expenditure ran past $350 billion in 2025, and Goldman Sachs projects $1.15 trillion across 2025 to 2027, more than double the $477 billion spent in the three years before. PwC models $31.6 trillion of data centre capex through 2050.

Somebody absorbs the volatility of bets that size, and right now it's the people working in and around tech. More than 245,000 tech workers lost their jobs in 2025, and 2026 has already passed 210,000 with AI the most-cited reason for cuts four months running. If you work in this industry, you already know the texture of it: the re-org you didn't see coming, the role that got redefined without you, the low hum telling you you're one tooling cycle from irrelevance.

Here's the uncomfortable part. The restructuring arrived years ahead of the returns. Companies reorganized around a capability that hasn't yet produced the outcomes used to justify the reorganization.

Now we can finally put a number on the time-suck

For years the cost of all this was a feeling. Two research programs have made it legible, and reading them together tells you far more than either on its own.

Dr. Rebecca Hinds at Glean's Work AI Institute surveyed 6,000 workers and found 6.4 hours a week disappearing into botsitting: feeding the model context, supervising output, debugging, cleaning up confident errors. Her framing is the sharpest version of it. If the productivity dividend gets spent on botsitting, you haven't removed work, you've added a layer of overhead.

Dr. Molly Sands at Atlassian's Teamwork Lab measures the other side of it. Her team puts 2.4 billion hours a year, inside the Fortune 500 alone, into hunting for information the company already has somewhere. Sands also found the upside is real but conditional: people who build AI across their workflows, rather than bolting it onto individual tasks, are recovering close to a full day every week.

Sit with the gap between those two findings, because it's the entire argument. Same technology, opposite outcomes, and the only variable is how the work was designed around it.

Here's what changed in 2026

What moved this year was the interface and the substrate, not the benchmark. OpenAI put ChatGPT Voice into the desktop app in July, where it directs Codex agents and drives computer-use tasks by speaking instead of typing. Dictating at conversational speed against agents running in the background collapses the setup cost that has kept most people from delegating anything that mattered.

Apple used WWDC to rebuild Siri on its third generation of foundation models, including a natively multimodal model that runs on device. The capability ceiling is low next to frontier models and the specs are tight. What matters is where it runs, because a private, secure, on-device tier now exists inside the ecosystem where you already keep your health data, messages, photos, and location.

Grok Bot, in beta since August, hands each user a persistent cloud computer with a browser, filesystem, and terminal, then organizes named agents under a chief of staff that routes tasks and escalates only when a decision is needed. Whether xAI wins doesn't matter much. The pattern does, and it's the first credible answer to botsitting: a durable environment, delegated specialists, and a human approving exceptions instead of supervising steps.

Developers have had a rough version of this for two years. It's now arriving for researchers, analysts, marketers, and operators, which is where the botsitting tax is being paid.

What this looks like when it finally works

Health is where the gap between data and outcome bothers me most. I've written before about how wearable baselines move detection from population averages to personal deviation. If data volume were the constraint, health outcomes across the OECD would already be climbing, and they are not. The constraint is inaction and a very high tolerance for feeling roughly fine.

What changes that is software that knows your baseline, notices the deviation, and has earned enough trust that you act on it. That's a behavioural design problem wearing a technical costume.

Education is the more urgent case, and the numbers landed this month. The OECD released PISA 2025 on September 8. Reading performance across member countries has fallen 28 points since 2015, roughly a year and a half of learning. Mathematics is down 22 points. One in five 15-year-olds is now a low performer across science, mathematics and reading, up from 16% in 2022.

I've spent most of my career working in and around education, and this is the part that keeps me up. The institutions that underwrote most of the last century's innovation are losing capability and relevance at the same time.

AI will either accelerate that decay or interrupt it. The version that interrupts it is unglamorous: helping a student articulate why higher education is worth pursuing at all, turning a vague interest into a direction, choosing courses and the experiences stacked around them against real career outcomes, and keeping a cohort connected through the years when most of them quietly drift away.

For the operators in between, the wins are just as boring and just as large. Surviving a tariff regime that rewrites itself quarterly, finding and qualifying for incentive and training programs, growing the capability of a small team: these are problems of synthesis and follow-through, not intelligence.

Every upside here comes with a bill

Every capability I've described arrives with a cost the market hasn't priced. The DOJ, FTC, UK Competition and Markets Authority and European Commission have jointly flagged concentrated control of specialised chips, compute, and technical expertise as a structural risk, and the concern has sharpened as long-term power arrangements let dominant firms foreclose energy capacity their rivals need. When a handful of companies own the substrate an economy runs its thinking on, they set the terms for everyone downstream.

Privacy pulls in the same direction, which is why Apple's constrained on-device models matter more than their benchmark position suggests. Local inference is the only answer anyone has offered that doesn't require handing a vendor the most sensitive data you hold.

Which brings me to design

AI has diminished design as a discipline, and pretending otherwise helps no one. When creation costs approach zero, the craft of making a thing stops being scarce. Curation is the next standard to collapse, as agents become the default arbiters of taste for people who no longer have the time or the trained judgment to arbitrate for themselves.

Which is why design is about to matter more than it has in a decade. No amount of model capability resolves what a person needs, what they'll tolerate, and what will move them to act. Turning something impressive into something useful is the entire remaining problem, and it's the one problem no frontier lab is organized to solve.

That work comes down to four barriers no model release clears on its own: inspired creativity, applied capabilities, technical complexity, and proactive relationships. Every AI product that sticks has crossed all four. Most of what's shipping right now has cleared the third and ignored the rest.

This is the work I do at PH1 Research. If your AI product demos beautifully and struggles to stick, scale, or justify its cost, the gap is almost never the model.

We spent four years proving this technology is impressive. The next four decide whether it was worth the bill, and that's a design question.

How helpful was this article?

Have a story to share?

0 / 500
A
Arpy Dragffy

Founder, PH1 Research · Co-host, Product Impact Podcast

Latest Episodes

All episodes

Product Impact Newsletter

AI product strategy delivered weekly. Free.