Introduction
Banks must identify the beneficial owners of every business customer, usually by tracing ownership through company registers. We built UBO-Bench to test whether current AI agents can do this reliably in production. We tested 13 models from Anthropic, OpenAI and Google in more than 5,000 runs across simple, complex and next-step cases from the UK and German registers. We also investigated what factors drive the agent performance. In particular, is the key ingredient the LLMs themselves or rather the country-specific rules guiding the agent?
Our main finding is that AI agents are ready for production in beneficial ownership checks as the leading models trace even the largest and most complex structures correctly. While smaller models, such as GPT-5.4 Mini and Gemini 3.5 Flash-Lite, lost parts of large structures, every model we compared scored at least 95% when configured with the complete register record, guidance on each country’s registers and clear criteria for when a case needs a person. About 90% of UK cases and 83% of German cases complete without needing more information, with the rest going to a person because the records leave a gap or a conflict that no register can settle. Rules help determine the next steps, such as a document to request or a conflict to resolve.
Key Findings
The leading models trace ownership correctly.
Four models, spanning Claude, GPT and Gemini, returned accurate results on all 55 real companies they were compared on, from two-node structures to structures of up to 28 companies and people or 12 levels deep.
Configuration matters more than the choice of model.
The frontier models are generally capable and country specific config and rules took all eleven compared models from between 75% and 90% to between 95% and 100%.
Most cases need no next step.
About 90% of UK cases and 83% of German cases complete without further information.
A person reviews only the cases the registers cannot settle.
Fixed rules send a case to a person when records disagree, information is missing, such as an unidentified controller or a German stand-in owner, or companies are registered as owning each other.
1. Background: beneficial owners and company registers
Before a bank opens an account for a business, it must identify the people who ultimately own or control it. This includes anyone holding more than 25% of the capital or votes, directly or through other companies, or controlling it by other means, such as the right to appoint the board. The EU’s anti-money-laundering regulation (AMLR) sets the threshold at 25% or more. In Germany, if no person qualifies as a beneficial owner, the law treats the managing directors as the beneficial owners, which we call stand-in owners.
In the EU this is not just a form the customer fills in. Ownership has to be traced through registries, layer by layer, until the chain reaches natural persons or a legitimate dead end. With the AMLR applying from July 2027, the topic sits at the top of most European compliance agendas.
Today an analyst does this by hand. Pull a register extract, read the shareholder list, find a corporate owner, pull that company’s extract, repeat. The task is inherently time-consuming and highly manual, but the main difficulty arises from two key factors: data quality and heterogeneity of information. Despite being the source of truth, registry data is known to be error-prone, yielding ambiguous situations due to spelling mistakes, contradictory information or plain mistakes in the data. Moreover, as unveiling ownership structures is a cross-border exercise, an analyst needs to work with the registries of various jurisdictions with different languages, legal forms and country-specific differences in the way they are structured.
This ambiguity and complexity make ownership structure unfolding in theory an optimal target for agentic automation, shifting the time spent by an analyst away from the tedious and error-prone data collection work towards reviewing agent output and applying human judgement where necessary.
2. How the agent works
The core design philosophy is to only apply agentic reasoning where needed while relying on rules and controlled data sources wherever possible. Ownership structure analysis is inherently composed of various deterministic steps, e.g. pulling data from the registry for each company encountered in the ownership chain. These steps are therefore handled deterministically while the LLM acts as the glue in the middle that resolves ambiguities, e.g. in case of misspellings, closed register sheets or country-specific nuances in the way register data behaves.
For our experiments the agent is invoked with just two inputs, the company name and the jurisdiction. It needs to source all other information from the registries. Note that in production this type of agent might also allow for additional inputs in the form of documents, free-form text as well as information sourcing via web research. These were deliberately deactivated to maximize comparability during testing.
To then actually enable a strong performance of the agent on the ownership resolution task, there are four key components in our agent design:
- Purpose built registry tools on a live connection. The agent searches the commercial registers, orders official reports and reads the shareholder lists and register extracts that come back. This is not implemented via free-form access to APIs and web pages but rather via specialized tools that normalize the way the registry data is requested and returned across jurisdictions. For the real companies in this study, live registry access was given via our data vendor partner Kausate, which enables live calls to corporate registries in various jurisdictions.
- Country specific guidance. Whenever the agent uncovers a new jurisdiction in an ownership chain, e.g. a German holding on top of a UK root company, it pulls a set of country-specific rules. These rules carry the nuances no general-purpose model knows reliably. For example, how a German seat transfer mints a new register number and closes the old sheet, how to handle circular ownership structures in German groups, and which legal forms publish no shareholder list so that a chain legitimately ends there.
- Graph building mechanics. The agent harness enforces a rigorous JSON schema on the agent output. This complex structure is defined as a set of nodes (companies and natural persons) as well as edges (ownership shares, control relations etc.), which the agent needs to underpin with actual sources via a mandatory evidence field. This schema also controls the additional information the agent needs to collect on each node, e.g. date of birth for natural persons or the officers of a company. Moreover, the schema allows the agent to report contradictory information, e.g. in the case of contradictory edges.
- Deterministic post processing. A rules engine applies the ownership logic (the 25% threshold, control chains, aggregation across paths) and raises a fixed set of “requires further investigation” flags. This allows the agent to focus on objective data collection while the interpretation, application of rules as well as the visual representation is offloaded to deterministic logic. Note that this is made possible by the strict schema enforcement on the agent output.
As mentioned above, the agent collects a plethora of information that goes well beyond the ownership structure itself. It includes actual source documents, registry identifiers, industry classification, officers, addresses and more. All of this is needed to utilize the agent output in an actual KYB process, i.e. to determine UBOs, screen for PEPs and sanctions or verify individuals within the chain. Most importantly, this can be transformed into an ownership graph.
Throughout this article, we explore elements of the whole setup including the model with its tools, guidance and output schema, and the deterministic rules on top. The model collects and connects the records and the rules calculate the beneficial owners and flag the next steps.
3. Test setup
The experiments for this article focus on two jurisdictions, the UK and Germany. We deactivated all other jurisdictions during testing to constrain the sampling space and avoid getting distracted by country-specific noise. This means that if the agent encountered a parent company outside of the geographical scope, its correct answer was to add this as a dead end to the ownership graph.
To collect our sample of companies to be tested, pure random sampling was not the right approach as a large majority of companies, in particular micro SMEs, have very simple ownership structures. However, the bulk of workload in KYB processes does not stem from these simple companies but rather from the more complex multi-layered ownership cases. To reflect this in sampling, UK companies were filtered from the Companies House bulk data for a single active corporate owner that resolves to natural persons. German companies were enumerated through our provider and screened two levels deep, because there is no bulk data to filter on. The simple and complex sets hold 104 UK and German cases with 40 simple companies, and 64 complex cases (including 15 large structures and 15 companies that each require a specific registry maneuver). A third set, the next-steps set, holds 121 cases in which the records may leave a gap or a conflict requiring some human next steps.
For the real companies, ground truth for evaluations was established by resolving the ownership structures using our production agent and then manually checking and adjudicating ambiguous samples by a human expert. During testing, a sample counts as ‘correct’ when its graph carries every company and natural person of the ground truth with correct linkage. Cosmetic differences to the ground truth, e.g. pooling minority shareholders vs displaying them in separate nodes, were manually reviewed and marked as correct if all the information was present.
For the model comparison and the next-steps set, all models got the same instructions, tools and output format, modeled on Taktile’s production ownership agent, and recorded an outcome for each company, such as owners traced, nothing to find, unresolved or outside the covered registers. Beneficial ownership can also come from control, such as a majority of the voting rights or the right to appoint the board, which public registry information alone does not show and our scoring does not cover. A case counts as correct only if the beneficial owners, every company’s outcome and the list of companies are all right, and any conflict between records, including companies registered as owning each other, is alerted as a requiring next steps from a human.
The 40 simple companies and the 15 large structures ran on seven models: Claude Sonnet 5, Claude Opus 5, Claude Haiku 4.5, GPT-6 Astra, GPT-5.4 Mini, Gemini 3.8 Flash, Gemini 3.5 Flash-Lite. Every model was run at the ‘low’ reasoning level, with the same guidance, tools and schema.
Lastly, note that token cost is not a major comparison dimension in our testing as actual agent run cost for ownership resolution is dominated by data costs. One shareholder report for a single company can cost more than $10 in certain jurisdictions and hence the token costs, which are far below $5 even for complex structures, are comparatively insignificant. Moreover, wall clock time is not reported as this is driven by registry responsiveness and not the LLM itself.
4. Tracing simple and complex structures
The first question was whether agents can do the core task of an ownership check: following the chain company by company until it reaches people. Our data includes large structures, up to 28 companies and people or 12 levels deep. The simple companies and the large structures break down as follows.
The simple set and the large structures of the complex set
| Set | Country | Companies | Nodes per graph | Ownership layers |
|---|---|---|---|---|
| Simple | United Kingdom | 20 | 3–18 (median 5) | 2–3 (median 2) |
| Simple | Germany | 20 | 2–16 (median 4) | 1–5 (median 2) |
| Complex | United Kingdom | 8 | 9–27 (median 11) | 3–12 (median 5) |
| Complex | Germany | 7 | 8–28 (median 18) | 4–8 (median 6) |
The examples below show three of the cases.
Correct graphs and model cost · seven models
| Model | Simple (40 companies) | Complex (15 companies) | Model cost per company |
|---|---|---|---|
| Claude Sonnet 5 | 100% | 100% | $0.32 |
| Claude Opus 5 | 100% | 100% | $1.02 |
| GPT-6 Astra | 100% | 100% | $1.20 |
| Gemini 3.8 Flash | 100% | 100% | $0.14 |
| Claude Haiku 4.5 | 98% | 93% | $0.25 |
| Gemini 3.5 Flash-Lite | 95% | 87% | $0.07 |
| GPT-5.4 Mini | 75% | 67% | $0.12 |
Correct means the graph carries every company and person of the ground truth, with correct linkage. Cost is the mean model cost per company on the simple set.
Four models, spanning Claude, GPT and Gemini across three price tiers, returned identical entity sets on all 40 simple companies, and the same four models remain at 100% on the 15 large structures. These results clearly show that for simple ownership structures the LLM model choice is insignificant. Complex structures require a model above a certain capability threshold, which all four leading models clear. The spread in model cost is 17x, from $0.07 to $1.20 per company, and in all instances token costs are insignificant compared to potential data costs.
A note on configuration: All models ran at the lowest reasoning level all seven support. Re-running only the failures one level up resolved 9 of 12, both of Flash-Lite’s and seven of GPT-5.4 Mini’s. While it may sound attractive to use cheaper models for simpler structures and higher tiers for more complex ones, this type of model routing is not possible in practice because the depth and complexity of ownership are not known a priori.
DetailWhy smaller models fall short on large structuresShowHide
On the simple set, Haiku 4.5 omitted the two trustee entries of a family trust on one UK company, Flash-Lite omitted one owner on each of two samples, and GPT-5.4 Mini omitted or added entities on ten and returned malformed graphs on eight. On the complex set, Haiku 4.5 omits an entire branch of a German group on one company, Flash-Lite loses a holding chain on one company and records the intermediate layers of another under names no other model found, and GPT-5.4 Mini fails five of fifteen and returns a defective graph on nine.
Entity loss and graph defects coincide. The models that omit entities are the models whose graphs are malformed: an edge to an undeclared node, a holding connected to nothing, two versions of a cap table kept at once. The failure is not in identifying owners but in carrying a 25-node structure through a multi-step procedure and emitting it as one valid object. The structured output described in section 2 becomes the binding constraint. The floor is procedural, not a gap in knowledge.
Every agent answer is emitted as a structured graph against a fixed JSON schema. A non-conforming output is rejected by the agent harness with validation errors and has to be re-emitted. This retry mechanic drives the token usage of lower-tier models, as the table below shows for the simple set.
| Model | Runs rejected by the schema | Rejected attempts | Tokens spent re-emitting |
|---|---|---|---|
| Claude Opus 5 | 0/35 (0%) | 0 | 0 |
| GPT-6 Astra | 0/33 (0%) | 0 | 0 |
| Claude Sonnet 5 | 3/36 (8%) | 3 | 18,191 |
| Gemini 3.8 Flash | 10/32 (31%) | 15 | 71,448 |
| Claude Haiku 4.5 | 19/34 (56%) | 51 | 261,313 |
| Gemini 3.5 Flash-Lite | 26/40 (65%) | 61 | 301,781 |
| GPT-5.4 Mini | 30/40 (75%) | 48 | 167,706 |
Rejection rates follow the tiers: the frontier models were never rejected, the middle tiers rarely, the small models routinely. Rejection does not by itself predict a wrong answer; Haiku 4.5 is rejected on more than half its runs and still finds the right entities on 39 of 40 simple companies. It is a cost paid in output tokens, and it erodes the price advantage that makes a small model attractive. On the complex structures, the graph defects are an artifact of multiple retries to emit a correct JSON structure. While the lower-tier models ultimately succeed, information is lost along the way.
Figures 4 and 5 show two situations that make tracing harder at any size.
Finding
Claude Sonnet 5, Claude Opus 5, GPT-6 Astra and Gemini 3.8 Flash traced 100% of simple and large complex structures correctly, while smaller models (like GPT-5.4 Mini) lost parts of large structures. Across eleven models, 98% of runs were fully correct on owners reached through two routes and 95% on names that may belong to one person or two.
5. Country-specific rules
The second question was whether that performance carries over to cases that rely on the nuances of a country’s specific register conventions. 35 cases in the dataset turn on a convention like this. For this experiment, we settle on one specific model: Claude Sonnet 5. This proved to be effective at both simple and complex tasks while still being significantly less expensive, by a factor of 3 to 4, compared to both Astra and Opus. The question we wanted to answer was, what happens when you turn country-specific guidance on or off? For this, we created a test set of 15 special cases (7 UK, 8 DE) with 2–22 nodes (median 6), which have some country-specific mechanics in the way their ownership structure behaves, e.g. a seat change in the German registry or a misspelling in the official registry data.
Claude Sonnet 5 · 15 cases with country-specific mechanics
| Guidance | Cases resolved correctly |
|---|---|
| Full country guidance | 15/15 (100%) |
| Generic guidance | 11/15 (73%) |
Without the country rules the agent loses 9 entities across 4 of 15 cases. This is one of the largest swings in the study. While Sonnet 5 cleared the capability bar in the previous section, it cannot overcome the country-specific quirks on its own despite the graphs actually being quite small.
These failures without country guidance are silent. The graphs are not visually broken but some pieces of information are missing which are required for UBO determination. These cases are also hard for human reviewers, so the model being able to solve them with country-specific rules is where an agentic ownership resolution becomes a massive value-add from a business perspective. The following is one such example:
ExampleA company that moves courtShowHide
When a German company moves to another court district, it is re-registered under a new number and its old entry is closed, but its parent’s filings may still cite the old number. An agent that does not know this finds a closed entry and stops, missing everyone above it (Figure 6).

Finding
Claude Sonnet 5 resolved 100% of the 15 cases correctly with country guidance and only 73% without, and eleven models given the guidance were fully correct in 98% of runs on moved registers. The guidance had no rule for heirs, and on those cases eleven models were fully correct in only 41% of runs; adding the rule took GPT-6 Astra and Claude Sonnet 5 from 0 of 20 runs to 20 of 20. Country-specific guidance is essential for achieving maximum performance, and it needs a rule for each convention the agent will meet.
6. Surfacing next steps for human review
Some cases need to raise next steps for a person to review. When a record is missing or two records conflict, the registers alone cannot complete the file. Instead, a person has to request a document, trace a parent abroad, or establish which record is right. This is true regardless of if a human or agent is doing the graph tracing. The agent records the gap that it found, and the rules on top turn it into a flagged next step. The next-steps set tests whether the whole setup surfaces the right cases to humans. It holds 121 cases, including cases with nothing missing (which the agent should resolve fully with no next steps), and cases where a next step should be identified for further human action. Three examples follow.
Example 1A shareholder list only on paper (Germany)Next step: request the paper listShowHide
A German shareholder list filed on paper before 2007 may never have been digitised, so the register shows no owners and a person has to request the paper list from the company (Figure 7).

Example 2An unidentified controller (UK)Next step: identify the controllerShowHide
A UK company files a PSC statement saying a person with significant control exists but has not been identified when it believes someone controls it but cannot name them. This means the beneficial ownership picture is incomplete, even when the shareholder list in the confirmation statement adds up to 100% (Figure 8). Of 28 active UK companies we checked with such a statement, 21 also had a full list naming someone above 25%.

Example 3Companies registered as owning each other (UK)Next step: find out which record is wrongShowHide
Some UK groups register a company as a shareholder of its own parent, which UK company law generally does not allow (Companies Act 2006, section 136). Tracing around the loop finds no one above 25%, so a rule flags the loop and a person establishes which record is wrong (Figure 9).

The general instructions tell the agent to trace each chain until it reaches people or cannot be continued. With them, every run surfaced a missing list, a conflict and a parent abroad, since each breaks a chain. An unidentified controller does not, because the shareholder list still reaches people. We therefore added a short paragraph of completion criteria: decide each company’s outcome from all of its records, not the percentages alone; keep it open if any record names an unidentified controller; and order each company’s report once.
Rules then run after the agent finishes and flag a next step if the agent marks a company unresolved or reports a conflict, if the UK register records an unidentified controller, if a German company names a stand-in owner, or if companies are registered as owning each other. Every other case is complete.
Right next step by situation and configuration
| Eleven models, 65 cases | Next step | General instructions | With completion criteria | Plus rules |
|---|---|---|---|---|
| Shareholder list only on paper | Request the list | 100% | 100% | 100% |
| Records that disagree | Resolve the conflict | 100% | 98% | 98% |
| Parent registered abroad | Trace the parent | 100% | 100% | 100% |
| Companies registered as owning each other | Determine the correct record | 35% | 30% | 100% |
| Unidentified controller (UK) | Identify the controller | 27% | 99% | 100% |
| Nothing missing | None | 99% | 99% | 99% |
On 20 unseen cases, half with an unidentified controller, the criteria raised correct handling of the controller from 25% to 95% of runs across the eleven models.
How much of the work completes on its own
The table lists each reason a case needs a next step and how often it occurs where we could measure it.
Reasons a case needs a next step
| Reason for a next step | How often | Flagged by our rules |
|---|---|---|
| Records disagree about ownership | About 9% of UK and 1% of German runs on live register data | Yes |
| The register says a person with control is unidentified (UK) | About 2,100 of 5.7 million active UK companies | Yes |
| No person qualifies, so the managing directors stand in (Germany) | About 15.7% of German limited-company filings | Yes |
| Companies registered as owning each other (UK) | 0.6% of UK chains with a corporate parent | Yes |
| A filing is missing or unreadable | Not measured on real data | Yes |
| A parent registered in a country not activated during the benchmark | 12.3% of UK chains with a corporate parent reach a parent outside the UK and Germany | No, recommended |
On live UK register data, agents found conflicting records in about 9% of checks and unidentified controllers are rare, so about 90% of UK checks need no next step. In Germany, conflicts came up in 1% of checks, but about 15.7% of limited-company filings name a stand-in owner (German federal government, answer to parliamentary question 20/11496, March 2024), so about 83% need none. In those cases, the file is complete and does not require a next step from a human. On the next-steps set, 99.5% of runs that completed without a next step were correct.
7. Model comparison
The leaderboard covers the 87 cases that all eleven models ran, 22 from the complex set and 65 from the next-steps set. A case counts as correct if the agent traced the right owners and surfaced the ones that need a next step. With the general instructions, the models scored 75% to 90%, while with the country specific config and rules, nine models scored 100% and none below 95%.
Complex and next-steps sets · two configurations
| Model | With country specific config and rules↓ | General instructions↕ |
|---|---|---|
| 1 | 100% | 90% |
| 1 | 100% | 83% |
| 1 | 100% | 82% |
| 1 | 100% | 82% |
| 1 | 100% | 82% |
| 1 | 100% | 81% |
| 1 | 100% | 80% |
| 1 | 100% | 80% |
| 1 | 100% | 80% |
| 10 | 99% | 79% |
| 11 | 97% | 75% |
The configuration closes the gaps between models. Under the general instructions, the models differed mainly on cases whose records looked complete, such as unidentified controllers and companies registered as owning each other. With the country specific config and deterministic rules, every model achieved perfect graph tracing and surfaced every case that needs a next step except one conflict missed by Gemini 3.5 Flash-Lite.
8. Implications
- Let the agent trace. The leading models are good enough, and on the simple set and the large complex structures Claude Sonnet 5 matched models costing three to four times as much.
- Give it country-specific guidance. This made the largest difference in section 5, from moved registers to shares held by heirs.
- Name the records that keep a case open, what must still be checked before stopping, and what each report costs.
- Supervise with deterministic rules. Let the agent collect and connect the records, and let rules calculate the beneficial owners and flag each case that needs a next step, such as unresolved and conflicting cases, UK statements of an unidentified controller, German stand-in owners and companies registered as owning each other.
9. Scope and limitations
The study covers companies in the UK and Germany and treats a parent registered elsewhere as the end of what the agent can trace. Every EU country keeps a company register with its own conventions which will each differ from the UK and Germany and require their own configurations. The US has no public ownership register and relies on the customer’s own certification of its owners (FinCEN customer due diligence rule, 31 CFR 1010.230), which requires a structurally different agent setup. The study only covers ownership from publicly available data. Non-public information, e.g. in form of veto rights and voting right agreements was not tested.
10. Conclusion
Today’s leading AI models are ready for production in beneficial ownership checks. An agent with purpose-built tools performs exceptionally well across the task of ownership resolution. Across the 40 simple companies and 15 large structures, from two-node structures to structures of up to 28 companies and people or 12 levels deep, the four strongest models all returned accurate results across the board. Claude Sonnet 5 turned out to be a well-rounded model choice for this task which clears the minimum requirements bar while also keeping costs under control.
Most importantly, the differentiating factor is configuration rather than the model. Country-specific expert guidance allows an ownership structure agent to also clear the cases that are an actual headache to human analysts, and given clear criteria for what keeps a case open, nine of eleven models were correct on every case in our comparison and none scored below 95%. About 90% of UK checks and 83% of German checks then need no next step involving human intervention. For the rest, the rules flag the next step, and a person resolves what a register cannot answer.
Taktile Labs, September 2026.