We gave an AI tool full control of a laptop. Pre-configured apps. Full internet access. The question was simple enough to fit on a napkin: can artificial intelligence actually do an office job?

Executives certainly think so. Over 200 tech companies slashed about 120,000 roles this year alone. Layoffs.fyi tracks the bleeding, and the names on that list—Meta, Oracle—cite AI as the primary driver. Cloudflare’s CEO didn’t mince words after firing 1,100 people. He expects AI to replace middle management, finance, and marketing.

We put those agents to the test. And the answer isn’t a clean “yes” or “no.” It’s a messy “sometimes.”

Testing AI Agents on Real Office Tasks

We deployed AI “agents” to act as office workers. The goal was to see if they could handle autonomous decision-making based on detailed instructions. We used benchmarks from Carnegie Mellon University and OpenAI to simulate real-world environments. These tests used fabricated spreadsheets, memos, and staff lists, along with carefully written prompts to define the AI’s role and available tools.

We ran these agents using Anthropic’s Claude Cowork app. This setup allowed the AI to take control of the computer. But control is one thing. Competence is another.

The agents excelled at tasks that involved writing computer programs. They struggled with everything else.

Can AI Handle Human Nuance and Interface Design?

Our first test involved gathering feedback from colleagues on Slack. We created bots to mimic coworkers, feeding the AI pre-written responses. The agent managed some curveballs without skipping a bat. When a name was misspelled—“Sara” instead of “Sarah”—it still messaged the correct person. When an attendee didn’t respond, it didn’t hang around waiting. It logged the employee with a zero score and moved on.

Simple stuff. But then we threw something complex at it.

The agent had to review documents and identify staff cuts for a government agency’s budget. It needed to decide which roles were expendable. The result? It produced a three-document report. Mostly accurate. It correctly identified that retirements and resignations could meet budget goals without firing anyone.

But it made a meaningful error. It flagged employees going on leave for layoffs. It lumped them all together, ignoring the fact that they planned to return.

Where a human would instinctively ask how long the leave lasted, the AI assumed they were gone for good. It inferred a permanent cut from a temporary absence.

“AI may be less capable of replacing tacit know-how, the idiosyncratic skills and tricks that accumulate with experience.” — Stanford and NBER

Tacit knowledge. Things you learn by doing, not by reading. AI doesn’t have that. Not yet.

Why AI Agents Prefer Coding to Clicking

The AI also choked on basic interface navigation. It struggled with the Google Drive interface. It fumbled through spreadsheets. When instructed to modify an organizational chart in Preview or Google Slits, it gave up immediately. It decided using an app was “too impractical” and wrote code instead.

Graham Neubig, a Carnegie Mellon professor and co-author of the AgentCompany project, noted that agents work in a “very unhuman way.” They prefer coding over typical user interfaces.

Our agent confirmed this bias. Despite explicit instructions to try the apps first, it ignored us. After about 30 seconds, it switched to coding. It recognized it was good at code. It recognized it was bad at clicking buttons designed for human eyes.

The next task reinforced this pattern. The agent had to fill out 17 I-employment verification forms using data from a spreadsheet. It uploaded them to Google Drive.

Again, it avoided the apps. It wrote code. The solution worked well. It even added formatting checks for phone numbers and Social Security IDs. It breezed through what would be tedious work for a human.

Then it stalled. A trivial step for any office worker: uploading the finished PDFs to the correct Google Drive folder. The agent tripped up. The code handled the logic, but the interface handling failed.

The Verdict: AI Needs a Human Boss

Our experiment matched broader findings in the industry. Scale AI recently tested agents on real freelance projects. The best-scoring model produced client-ready results only 16 percent of the time.

AI can excel at specific, isolated tasks. It can handle structured data. It can write code. But it fails at the nuanced, intuitive, and interface-heavy parts of most jobs. It doesn’t understand the office politics. It doesn’t know when someone is just taking a break versus when they’ve quit.

It can add value to certain areas of the workforce. But for now, AI still needs a human to watch its back. A human boss. Someone to catch the uploads that got stuck. Someone to correct the hiring mistakes.

The tools are getting better. But the gap between “doing a task” and “doing a job” is still wide.