Skip to content
Breaking
Latest technical intelligence from Northeast India • Infrastructure, AI, Cloud & Security Analysis • Precision Analysis | Raw Intelligence | Your North Star of Tech Latest technical intelligence from Northeast India • Infrastructure, AI, Cloud & Security Analysis • Precision Analysis | Raw Intelligence | Your North Star of Tech
TECHNOLOGY

Analysis: New study shows AI isnt ready for office work

AI's Office Test: A Reality Check for the Knowledge Work Revolution

AI's Office Test: A Reality Check for the Knowledge Work Revolution

In a bold prediction two years ago, Microsoft CEO Satya Nadella foretold the impending takeover of knowledge work by generative AI. However, a visit to a typical law firm or investment bank today would reveal a human workforce still firmly in control. A new study from training-data company Mercor sheds light on why the anticipated robot revolution is yet to materialize.

The Messiness of Real Work

The study, titled APEX-Agents, offers a stark reality check. Unlike traditional tests that ask AI to write a poem or solve a math problem, this one uses real-world queries from professionals. The models are expected to complete multi-step tasks that demand jumping between various types of information.

The Limits of AI's Capabilities

Even the top models on the market, such as Gemini 3 Flash and GPT-5.2, struggled to crack a 25% accuracy rate. Gemini led the pack at 24%, with GPT-5.2 close behind at 23%. Most others were stuck in the teens.

The Role of Context

Mercor CEO Brendan Foody underscores that the problem isn't intelligence; it's context. In the real world, answers aren't served on a silver platter. A lawyer, for instance, must check a Slack thread, read a PDF policy, look at a spreadsheet, and then synthesize all that to answer a question about GDPR compliance.

AI: The Unreliable Intern

Humans do this context-switching naturally. AI, however, is terrible at it. When forced to hunt for information across scattered sources, these models either get confused, provide the wrong answer, or simply give up.

Implications for Job Security

For those concerned about job security, this study offers some reassurance. The results suggest that currently, AI functions less like a seasoned professional and more like an unreliable intern who gets things right about a quarter of the time.

The Rapid Progress of AI

While this might offer some comfort, the progress is terrifyingly fast. Foody noted that just a year ago, these models were scoring between 5% and 10%. Now they are hitting 24%. So, while they aren't ready to take the wheel yet, they are learning to drive much faster than we expected.

Reflections and Looking Forward

As we navigate this digital transformation, it's crucial to understand the current limitations of AI. In the context of North East India and the broader Indian landscape, this study underscores the importance of human expertise in knowledge-intensive fields. However, it also serves as a reminder to stay vigilant and adaptable in the face of rapid technological advancements.