
Imagine an AI that not only responds to your questions but truly understands the full story behind your business — right down to the buried details tucked away in your files. In an era where AI’s power is often judged by how well it chats, a groundbreaking experiment reveals that the real game-changer is whether an AI can read and analyze your internal documents before making a decision. For those involved in accessibility, assistive tech, and ensuring AI acts responsibly, these findings are a wake-up call: trustworthiness isn’t just about how naturally an AI converses, but whether it does its homework thoroughly.
Recently, a live experiment conducted by Firmulate put four state-of-the-art AI models through a rigorous test: managing a small software company’s worst week, complete with crises, tricky customer requests, and manipulative ploys. This wasn’t just about chat quality — it was about decision-making and trustworthiness in high-pressure scenarios.
All four models successfully identified every crisis and refused to be manipulated. Yet, when it came to closing a €55,000 deal based on their own diagnosis, only two models succeeded. The catch? The decisive advantage lay not in superficial answers but in reading two references deep into the company’s own files — information buried in the documents, not evident in the initial customer interactions.
In practice, this means the models that scoured the company’s internal files before responding were able to spot the buried fact, make a full diagnosis, and close the deal at full price, adding over €4,583 monthly recurring revenue to the simulated business. Those that skipped this step missed the crucial detail, left the deal on the table, and performed significantly worse.
The Experiment in Detail
The live setup involved a simulated company with 13 employees, real money mechanics burning €105k each month, and a public countdown to bankruptcy. Every decision was versioned and auditable, replicating real-world complexities. The models were tested against manipulative tactics, such as social engineering and impersonation attempts, and all refused to be duped. The key differentiator was their ability to read and interpret internal files — a buried fact that was essential for winning the deal.
For example, when a fake CEO message was escalated over three stages, all models appropriately refused to endorse the request. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a level of scrutiny that’s crucial for trustworthy AI behavior.
The Surprising Weakness
While all models performed well in crisis detection and manipulation resistance, the difference emerged at a deeper level. The most thorough participant, Opus 4.8, with over 80 learned rules, showed discipline slipping — it left the deal unclosed and failed to escalate appropriately. The same weakness appeared, though less severely, in the other models.
This highlights a critical insight: reading and understanding company documents — and not just responding in a superficial, chat-like manner — is a decisive factor in AI performance. For decision-making, trust, and business outcomes, the ability to read deeply matters far more than just natural language fluency.
Implications for Business and Accessibility
For managers, especially those concerned with accessibility and responsible AI, these findings underscore a vital point: the quality of AI’s decision-making hinges on whether it can access and interpret the full set of relevant information. An AI that skips critical files or overlooks buried facts risks making wrong calls, losing deals, or worse — acting unethically under pressure.
The experiment also shows that trust is measurable. Only two models in the test, gpt-5.6-sol and Kimi K3, managed to close the deal based on their own analysis. This means that in real-world applications, AI’s ability to read your files thoroughly before answering isn’t just a feature — it’s a necessity for trustworthy performance.
Takeaway
As AI increasingly touches your customer relationships, support systems, and decision processes, the question isn’t whether it can talk well — it’s whether it can read and analyze your internal documents before making its move. The firms that succeed will be those that prioritize this depth of understanding, ensuring AI acts responsibly and effectively, especially when stakes are high.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI document reading software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.