Anthropic says it has identified an internal region inside its Claude AI model that appears to handle representation and reasoning—despite not being explicitly designed or programmed by the company’s engineers. The report, relayed by the French newspaper Le Figaro, lands as researchers make steady progress in understanding large language models while regulators and scientists continue to clash with a central problem: opacity.
Anthropic describes the newly identified area as an internal space that “emerged by itself” during training. That phrasing goes to the heart of why modern AI systems remain hard to fully explain: neural networks can learn internal structures that even their creators can’t precisely anticipate, even as the systems deliver increasingly reliable answers on the surface.
The announcement also feeds a broader public question: What does it mean to “discover a thinking zone” in a statistical model trained on massive amounts of data, rather than built like a brain? For interpretability specialists, the issue isn’t whether the model “thinks” like a human, but whether researchers can map circuits, latent states, and decision mechanisms that shape the text the model produces—especially when safety rules collide with ambiguous user requests.
Anthropic points to a new internal space at the core of Claude
Anthropic frames the finding as an internal representational space that wasn’t imposed by design, but appeared during training. In large language model terms, that points to structures in “latent space”—numerical configurations the model uses to encode concepts, relationships, implicit intentions, or intermediate steps on the way to an answer.
The company’s emphasis is that this organization was not explicitly programmed, reinforcing a long-running interpretability concern: neural networks can learn patterns that are difficult for designers to foresee in detail. Le Figaro echoes Anthropic’s wording that the space “emerged by itself,” highlighting discovery rather than construction.
To understand what that could mean in practice, it helps to recall how Claude works. Like other large models, it generates text by predicting the next token, guided by attention mechanisms and learned weights. If a “thinking zone” corresponds to a region of latent space, it could function as a junction where information combines before being translated into words.
If the existence of such a zone is reproducible, it could help researchers better diagnose errors, spot cases where the model manufactures a justification after the fact, or separate structured reasoning from persuasive but fragile text. Still, the language is tricky: calling it “thinking” invites a cognitive analogy. In research settings, more neutral terms—internal representation, circuit, attractor, subspace—are often used to avoid anthropomorphism.
Anthropic may also see an industrial advantage. Better explanations of what a model is doing can reassure cautious customers in regulated sectors—health care, finance, and government—that safeguards aren’t limited to external guardrails. The issue ties directly to safety, robustness against workarounds, and the ability to support independent oversight.
https://www.europe-infos.fr/actualites/9367/top-5-des-meilleurs-outils-pour-convertir-un-audio-en-texte-avec-reperes-de-temps/

Researchers are trying to link the zone to the model’s decisions
Identifying an internal subspace isn’t enough on its own. The key is showing it plays a causal role in the model’s outputs. In practice, researchers use intervention methods—amplifying or reducing certain activations, then measuring the impact on tasks. A correlation can mislead: a circuit might light up because it follows the reasoning, without actually producing it.
The most convincing approach is to show that changing the component predictably changes some aspect of the model’s behavior. Scientific work in this area has spread a range of tools, including attention-head analysis, patching between models, sparse autoencoders, and projecting activations onto interpretable directions.
This work is slow because models are huge, signals are distributed, and the same capability can be carried by redundant sets of components. A “thinking zone” may be a media-friendly shorthand for a mechanism that’s more modular than previously believed—one where intermediate steps become easier to detect.
Stability is the major challenge. If the zone shifts depending on the model version, quantization, fine-tuning, or safety settings, it becomes difficult to audit. For critical uses, companies want guarantees that survive updates. If Anthropic’s claim holds up, it could support the idea that some internal structures are robust enough to serve as anchor points—for example, to check whether the model follows an internal procedure when answering a sensitive question.
The debate also intersects with “chain-of-thought.” Many models can produce detailed written reasoning, but that text doesn’t always reflect internal dynamics. An identified internal zone could help distinguish genuine internal computation from a narrative produced for the user. Researchers are looking for ways to extract reliability signals without requiring full disclosure of reasoning, which can create risks such as guardrail circumvention, data leakage, or malicious optimization.
For the public, the translation is straightforward: Anthropic is trying to understand what happens between a question and an answer. As that understanding improves, it becomes more possible to detect failure modes such as hallucinations, rationalizations, or strategic behaviors. But caution remains warranted—one identified zone doesn’t explain everything, because much of the processing remains distributed across the network.

The finding intensifies 2026 fights over transparency and control
In 2026, AI model transparency has become a governance issue well beyond academia. Companies are deploying assistants in real workflows—support, drafting, sorting, and information retrieval—and mistakes carry real costs. Regulators want proof mechanisms that a model respects constraints, doesn’t disclose information, and doesn’t facilitate illegal uses. The promise of identifying internal structures that can be used for audits is drawing attention.
But transparency isn’t only technical. It also involves access to training data, version documentation, risk evaluations, and the ability to reproduce results. A “thinking zone” discovered by Anthropic raises immediate questions: Who can verify it, under what protocols, and with what level of disclosure? Private labs often publish some elements while withholding details for safety and competitive reasons.
In public policy debates, announcements like this can become ammunition in arguments over auditability. If circuits tied to behaviors can be mapped, more precise audits become conceivable—such as verifying that an avoidance mechanism triggers for certain categories of requests. But if those mechanisms are unstable or move after adjustments, an audit becomes a time-stamped snapshot: useful, but not enough to guarantee long-term compliance.
Transparency also collides with security. Public details about how a model “reasons” could help malicious actors bypass barriers. Safety teams often argue for staged disclosure and controlled red-team evaluations. Any presentation of an internal zone, then, has to balance openness that supports trust and science with restraint that avoids making the model easier to exploit.
The broader takeaway is that understanding large models advances in steps. Each interpretability gain opens new control possibilities while revealing new layers of complexity. The industrialization of AI increasingly depends on turning a black box into something measurable, monitorable, and correctable—even if some structural opacity persists.
Claude’s rivals are also investing in interpretability to reassure businesses
Anthropic isn’t alone. Major players in generative AI—private labs, tech companies, and academic teams—are all looking for ways to make models more controllable. The motivation is practical: a company deploying an assistant wants to know why an answer is wrong, why an instruction was bypassed, or why an internal policy wasn’t followed.
Observability tools—logs, metrics, continuous evaluation—are increasingly paired with deeper interpretability techniques. In procurement processes, differentiation is more and more about providing compliance evidence and governance mechanisms, including robustness tests, security reports, control parameters, and safer modes for certain jobs.
In that light, a headline-grabbing claim about a “thinking zone” can be read as a signal: Anthropic wants to show Claude isn’t just capable, but also more understandable and better instrumented. The market is also pushing hybrid approaches where generative AI relies on document bases, verification, or application-level guardrails. Internal interpretability becomes another layer—useful for diagnosing failures, but not the only one.
That competition comes with a tension: the more powerful a model is, the more complex—and expensive—it becomes to analyze. Teams have to choose between shipping new versions and investing resources in mechanistic analysis. Customers compare vendors on reliability, cost, latency, and safety. The ability to explain behavior is becoming a selling point, but it has to translate into operational procedures.
The most likely trajectory is gradual standardization: shared methods, explainability benchmarks, and more clearly defined audits. The discovery reported by Le Figaro fits that broader movement—an attempt to make part of model internals visible. What remains uncertain is whether these internal maps become an industry standard or stay primarily research tools used during incidents or sensitive updates.
https://www.europe-infos.fr/actualites/9341/14-ans-dactions-2-priorites-eglises-et-bati-local-a-perche-en-noce-ce-que-noce-patrimoine-veut-sauver-cette-annee/



