When I bought my house, one of the inspectors sold us on a roof inspection.
That seemed like a sensible thing for a new homeowner to buy. We needed help understanding the house. An inspection was supposed to give us information we could use to make a decision.
The inspection turned out to be a sales system. The inspection team also had roofing solutions to sell, and the process moved us toward buying those solutions. Helping a new homeowner understand the roof had become secondary to creating another sale.
The problem was the purpose the inspection served.
That experience is a useful place to begin with AI evaluation. When someone hands you an assessment, whose decision is it designed to improve? What happens when the evidence points toward buying less, delaying the work or walking away?
Those questions became more concrete on Sept. 18, when Anthropic announced “Partnering with Accenture on embedded evaluation.” Faculty, Accenture’s specialist AI business, will lead the work, with Anthropic funding it directly. The stated scope includes evaluating models, examining alignment and testing safeguards.
When Evaluation Meets an Existing Sales Relationship
There is an existing commercial relationship behind that announcement. In December 2025, Accenture and Anthropic announced a partnership that included a dedicated Claude practice and joint offerings to help enterprises adopt AI. That earlier announcement is useful context: the organizations already had a business relationship oriented toward expanding deployment.
Now there is an evaluation relationship too.
That combination deserves examination. It does not establish that Accenture or Faculty has compromised an assessment, suppressed a finding or failed to act independently. My roof inspection supplies a question to investigate. It supplies no evidence about their conduct.
The question is whether an evaluator can deliver a commercially inconvenient judgment, and whether the people who need that judgment will receive it.
Independence Has to Survive Ordinary Pressures
In “Dev, Test, and Prod Still Matter: What Gets Deployed Has Changed,” I argued that independent evaluation requires attention to the machinery around the test: access, isolation, evidence retention and responsibility when something goes wrong.
There is another layer around that machinery. People choose the questions, negotiate the scope, assign the staff, renew the contract and decide where the report goes.
A technically excellent test can sit inside a relationship that makes some questions difficult to ask. An evaluator might have access to the model while lacking a protected route for escalating a disagreement. A customer might hear that an evaluation occurred without learning which important questions it left unanswered.
These are possibilities to investigate, not findings about the new partnership. They explain why the commercial arrangement belongs in the assessment.
In “Winning the Intelligence Inversion,” I described independent risk governance as part of the operating model for AI. This development brings that principle into a purchasing decision. Independence has to survive the ordinary pressures of an organization: budgets, reporting lines, important accounts and the desire to keep work moving.
Nobody has to be dishonest for those pressures to matter.
Suppose an evaluation raises a problem that would delay a major rollout. The immediate issue is technical. The consequences reach account teams, delivery commitments, customer expectations and revenue. A credible arrangement should already explain who decides what happens next. Leaving that decision until everyone is under pressure makes the evaluator’s job harder.
The Case for Embedded Evaluation
But there is a strong argument for the arrangement Anthropic has announced.
People who help organizations deploy AI may understand failures that a distant reviewer misses. They may recognize the difference between a convincing demonstration and a workflow that breaks when permissions change, records conflict or a user does something unexpected. Embedded access could also reveal decisions and practices that cannot be reconstructed from a finished model.
Distance can protect judgment. It can also limit understanding.
Will Knight’s Sept. 18 WIRED article, “Here’s How an AI Slowdown Could Actually Be Enforced,” captures this disagreement. Geoffrey Irving argues that inspections can help constrain development; Raymond Douglas emphasizes unresolved questions about effective controls. The reporting offers reasons to investigate inspection methods carefully, alongside a credible case for using them.
Payment alone cannot settle the question either. Evaluation takes skilled people, time and access. Someone must fund it. Moving the invoice to a nonprofit, a pooled fund or a government body changes the incentives without eliminating the need to examine them.
Anthropic acknowledges that access, reporting and funding standards remain unsettled. It describes a nonexclusive arrangement, plans for additional evaluators and discussions about other funding approaches. Those are relevant commitments. Their practical value will depend on how they operate.
What Protects an Uncomfortable Finding?
There are also proposed protections that deserve a fair hearing.
In his September essay, “We Must Pace the Frontier,” Dario Amodei proposes allowing embedded reviewers to publish key findings without Anthropic’s editorial control. He describes limited redactions for specified sensitive information and says reviewers should be able to disclose when a redaction materially affects their conclusions.
Those proposals address a central concern in this article. They would give an evaluator a way to make an uncomfortable conclusion visible.
They should also be described accurately. A proposed protection is different from a verified provision in a particular signed agreement. The public partnership announcement does not spell out appointment and removal protections, commercial separation or the complete publication arrangement. That leaves questions for reporting. It does not prove those safeguards are absent.
Disclosure itself needs judgment.
A report could contain customer information, exploitable vulnerabilities or other material that should not be published openly. Demanding unrestricted publication would create problems of its own. Buyers should instead ask who can receive detailed findings through protected channels, how redactions are reviewed, how disagreements are escalated and whether a public summary can explain the assessment’s limits.
The useful distinction is between protecting sensitive information and preventing an adverse conclusion from reaching someone who can act on it.
A Model Assessment Is Not a Deployment Decision
There is a further limit to the roof-inspection comparison. In my case, we were the customers buying the inspection. An enterprise relying on a model developer’s evaluation may have no direct contractual relationship with the evaluator. It may receive a summary intended for a different audience and a different decision.
That makes scope especially important.
An assessment of a model developer’s safeguards can be valuable while leaving many questions about a customer’s deployment unanswered. Which records does the application retrieve? Whose permissions does it use? What can it change? Can an affected person challenge the result? Does the system behave acceptably in the languages and operating conditions where the organization uses it?
A model-level assessment cannot answer those questions simply by existing.
The Electronic Frontier Foundation supplies a useful challenge to a narrow view of safety. Its Sept. 18 “Statement on California Governor’s Executive Order on AI” supports third-party investigations while calling attention to present harms involving employment, benefits, surveillance and pricing. EFF is an advocacy organization; its statement identifies priorities rather than proving the effectiveness of a particular inspection method.
The distinction matters for an enterprise reader. A bank, an employer and a public agency can use the same model and create quite different consequences. Their customers and employees encounter the whole service, including the policies and decisions wrapped around the model.
In “How Much Work Can Your AI Safely Own?” I separated capability, evidence and authority. An external assessment can contribute evidence. The organization still has to decide what work to authorize and remain responsible for that decision. Anthropic likewise states that embedded evaluation does not transfer its responsibility for model safety.
Five Questions for Enterprise Buyers
So what should an enterprise buyer actually request?
Start with the decision the evidence is supposed to support. Expanding an internal drafting assistant and allowing an agent to change customer records require different cases. Then ask for enough information to understand five things:
- Scope: What system, version, configuration and use cases were examined? Which relevant questions were excluded?
- Access: What could the evaluator inspect, and what limitations affected its conclusions?
- Commercial safeguards: Who appoints, pays and can remove the evaluator? How are related implementation interests disclosed and managed?
- Reporting: Who receives adverse findings, who can redact them and where can an unresolved disagreement go?
- Follow-through: Who owns remediation, what evidence demonstrates improvement and what changes trigger another assessment?
These are proposed purchasing questions, not claims about a universal standard already in place. Their depth should match the consequence of the work.
Useful Scrutiny Has to Be Affordable
There is a cost argument here that should not be waved away. A medium-sized organization cannot necessarily commission a bespoke investigation of every supplier. An elaborate assurance requirement could favor established providers and leave smaller buyers dependent on reports they have little ability to question. EFF’s call to make investigations accessible to smaller developers points toward the same concern.
Reusable evidence, common reporting formats and proportionate reviews could help. A shared assessment should make its boundaries easy to understand so each customer can concentrate on the remaining questions in its own deployment. Repeating the same expensive inspection everywhere is not the only way to obtain useful scrutiny.
Nor is waiting for perfect independence a satisfactory plan. It could leave organizations relying on less evidence while institutions and methods develop. The practical task is to improve the arrangement, expose its limitations and keep testing whether it works.
Keep the Option to Say No
The evidence could ultimately support the Anthropic–Accenture approach. Protected disclosure, meaningful separation of commercial interests, sufficient access and a record of consequential findings would make a stronger case for embedded evaluation. Deployment expertise could prove to be one of its advantages.
That possibility belongs in the argument. The purpose of scrutiny is to learn whether an arrangement deserves trust.
The roof inspection failed us because the process served the next roofing sale more than it served a new homeowner. We needed help making a decision about our house. The inspection team had built a way to sell us more solutions.
An enterprise needs an assessment that leaves several choices open: proceed, narrow the deployment, repair the problem, obtain more evidence or decline the purchase.
A useful evaluator must be able to support any of those choices. The people making the decision must be able to see why.
Before accepting the next AI assessment, ask what happens if its most useful conclusion would reduce the next sale.
That was the question inside our roof inspection. It belongs inside the AI contract too.