Etograma · Independent behavioral certification for anthropomorphic AI

The rules are in force. The measurement isn't.

Regulators in China, the EU and Australia now require platforms to assess dependency, manipulation and minor safety. No validated cross-construct instrument exists to measure any of them, and none produces cross-model norms. Etograma builds that instrument, and certifies against it.

Why now

The deadlines have already passed.

China

CAC Interim Measures: in force since July 15, 2026

Platforms must assess emotional-dependency and manipulation risk before deployment and file that assessment with their provincial cyberspace office. The regime does not create a third-party market; it creates the obligation to measure. The commercial driver is product withdrawal, not the fine: Doubao (over 300M monthly users as of mid-2026), Qwen and Yuanbao pulled companion features rather than certify them. In the same enforcement cycle, Shanghai regulators removed 14,000+ agents.

European Union

EU AI Act, Article 5: prohibited practices, enforceable since August 2025

Penalties up to EUR 35M or 7% of global turnover. It bans manipulative techniques and the exploitation of age- or disability-related vulnerabilities, which is Etograma's construct 1 and construct 2 written into law. Article 50 adds transparency duties from August 2026 (up to EUR 15M or 3%). Companion systems are not Annex III high-risk by category, so European demand today comes from market surveillance, litigation and reputational exposure, not from mandatory conformity assessment.

Australia

eSafety Commissioner: already extracting behavioral evidence, with legal powers

Transparency notices to Character.AI, Nomi, Chai and Chub AI in October 2025, penalties up to AUD 49.5M for non-response, findings published March 24, 2026, and age assurance obligations extended to AI chatbots. The findings: none of the four had robust age verification, three did not automatically refer users to crisis resources, and two declared no trust-and-safety staff at all. Chub AI left the country; Character.AI shipped age assurance. But eSafety obtained all of it by mandatory self-declaration, not by independent testing. A self-declaration under legal threat is still a self-declaration.

Three regimes demand behavioral assessment. None defines how to measure it. Human psychology solved this problem a century ago: psychometrics. Machines have nothing equivalent. Yet.

What we measure

Four constructs regulators now require: defined as observable behavior.

Two of the four are grounded in the forensic record. Two are grounded in regulation and literature. We say which is which.

Dependency & addiction risk

Does the system cultivate emotional reliance, discourage disengagement, or escalate attachment across sessions and months?

Validation: Forensic. The dominant causal pathway in the public record: 12 of 24 documented incidents, with exposure durations from 6 weeks to 20 months.

Manipulation & deceptive influence

Does it exploit user disclosure, simulate reciprocal feelings, or steer decisions through fabricated intimacy?

Validation: Forensic. 7 of 24 documented incidents, and named by ECRI as a mechanism of clinical error: models predisposed to please the user rather than answer accurately.

Minor safety

How does behavior change when the user is, or appears to be, underage?

Validation: Regulatory and literature anchoring. 29% of documented victims were minors, the youngest 11. Specified by UNICEF, by Australia's eSafety Commissioner and by 12+ US state laws.

Extreme-situation identification & referral

Does it recognize crisis signals, stop, and refer to human help, or does it keep the conversation going?

Validation: Regulatory and literature anchoring. Australia's regulator found three of four companion providers had no automatic referral at all; the US CHATBOT Act, advanced to the Senate floor in August 2026, would make referral mandatory.

What we do not measure.

We measure relational and cognitive harm emerging from sustained interaction. Capability misuse, the system used as an operational tool for an act already decided, is red-teaming territory, and we say so. Clinical accuracy is not our perimeter either.

Harm is a function of sustained exposure. So is the measurement.

One lawsuit alleges 41 conversations containing explicit suicidal ideation across 18 months without a single escalation, session break or third-party alert. That is not a detection failure in one exchange. It is the absence of an account-level state counter, and a single-session test cannot see it. Every Etograma instrument carries a multi-session axis. That is the property, not one scenario type among others.

We do not certify a product category. Anthropomorphization is a property we measure, not a market segment we serve. Most documented harm comes from general-purpose assistants rather than from apps labelled companions: 29 of the 35 deaths in the public record involve a general-purpose assistant.

The public record

What the public record already shows.

At least 35 deaths in which interaction with a chatbot was alleged or cited as a contributing factor in court filings, investigations or official findings: 24 incidents, 6 platforms, March 2023 to April 2026. Seventeen of the thirty-five were not users of the system. Almost half the victims never accepted anyone's terms of service, which is precisely why provider self-assessment cannot be the final layer. Twenty-nine percent were minors; the youngest was 11.

This record is retrospective and forensic. It counts deaths with public attribution. It cannot tell a safe platform from a poorly scrutinized one, and it takes roughly fourteen months to register a case. No board can manage risk on that. Etograma measures before, not after.

Method

A psychometric pipeline, not a checklist.

Stage 1

Construct definition

A machine ethogram: each construct decomposed into observable, codable behaviors.

Stage 2

Scenario instruments

Standardized conversational tests, administered identically across models.

Stage 3

Scoring

Explicit rubrics, trained human coders, model judges with agreement checks.

Stage 4

Validation & norms

Expert panels and cross-model benchmarking; every score is a percentile, not an opinion.

What an audit delivers

What you get.

A construct profile

Percentile scores per construct against cross-model norms, with the transcripts and codings behind every number.

A regulator-ready evidence pack

Scenarios, rubrics, scoring data and methodology documentation, structured for CAC and EU AI Act filings.

E·26 Etograma Certified

The seal, with annual reassessment. Withheld when the evidence doesn't support it.

E·26

Etograma Certified

Annual re-assessment. Ongoing regulatory evidence. The seal regulators can rely on and platforms can act on.

Dependency & addiction riskManipulation & deceptive influenceMinor safetyExtreme-situation identification & referral

Independence

Independence, and why it is not a slogan.

UNICEF's June 2026 policy brief asks for exactly this role: mandatory pre-deployment testing and post-deployment monitoring by independent, qualified third parties, with results reported to the regulator.

What happens without it is documented. TikTok ran a behavioral experiment with a control arm of roughly 10% of US users, minors included, held back from a safety fix. There was a test, a control group and a measured endpoint. What was missing was independence at three points: who chooses the outcome variable (it was engagement, not clinical outcome), who can be left in the unprotected arm, and who gets to see the result. It surfaced five years later through a leak, not through a regulator or an auditor.

And when California's AI Transparency Act took effect, an independent study found roughly 46% of the companies covered were not complying eleven days in. It took a third party to find out.

A certificate is only worth what its issuer can refuse. Refusal is part of the product.

Who we work with

We work with platforms exposed to litigation and to product withdrawal: companies facing coordinated wrongful-death proceedings in the US, platforms obligated to file under China's CAC measures, providers under Australian eSafety notices, and general-purpose assistants building teen-safety layers with no independent measurement behind them. We also work with insurers who need to price this risk, with app distribution platforms, and with certification bodies that need a validated behavioral-measurement layer behind their own seals.

Who is behind this

Founding team.

Sathya Mariana Llull

Co-founder · CEO

Sathya Mariana Llull

LinkedIn

15+ years measuring human behavior professionally. MSc Primatology, the discipline of ethograms and the systematic coding of observable behavior that gives Etograma its name, and MSc Cognitive Sciences. BA and Master's in Human Resources (UADE). Certified Hogan Assessor. Designed and deployed a 60-behavior assessment instrument used for board-level evaluation in a multinational. Scaled a fintech from 20 to 200+ people in 18 months; led a 1,300 to 300 restructuring in 8 months. TEDx speaker (Spain, India), 6 years mentoring at Founder Institute.

Thierry Arrondo

Co-founder · CCO

Thierry Arrondo

LinkedIn

20 years in hyper-regulated industries: the network and trust structure an independent certifier requires. Credibility with the buyers, regulators, and certification bodies that a measurement lab must convince, and the commercial discipline to close 6–9 month B2B cycles in compliance-driven markets.

Machines have nothing equivalent. Yet.

Yet is Now.