Context Layer for Odoo
Benchmark methodology.
How the context layer is measured: a three-arm evaluation that holds the model, connector, permissions, and tasks constant and changes only whether OCL context is present. The design is public; independent blind review is pending and results are forthcoming.
The question a benchmark has to answer is commercial, not cosmetic:
Does the context layer cause materially more correct and safer Odoo outcomes than a strong connector plus raw schema, at an acceptable context, latency, and cost?
It is not a prompt beauty contest and not a set of screenshots. This page describes the method.
No results published yet
Live model outputs have been generated. Independent blind human scoring is the final evidence gate, and it is pending. Until that review is complete and passing, we publish no accuracy percentage, no chart values, and no claim that a gate has passed. What follows is methodology.
Three arms
Every arm uses the same model, connector capabilities, user permissions, tenant fixture, tool budget, and task wording. Only the context differs.
| Arm | What the model gets |
|---|---|
| A | Connector only: tool descriptions and returned records, no schema dump, no OCL. |
| B | Connector plus the strongest realistic raw schema and concise generic Odoo guidance. This is the primary baseline. |
| C | Exactly B, plus an assembled OCL context pack. No extra tools or permissions over B. |
B versus C is the comparison that matters. The connector, tools, permissions, and model are identical across those two arms, so any difference is attributable to the context layer alone and nothing else. Arm A is a floor; optional diagnostic arms (generic documentation RAG, full ontology dumping) may run, but claims center on the three stable arms.
What each case measures
Tasks are scored across levels, because an eloquent wrong answer is still wrong:
- Interpretation: the correct business noun, model, discriminating domain, time field, state, company and currency considerations.
- Query plan: valid fields, joins, aggregation grain, signs, and filters, without silently changing the question.
- Executed result: the plan runs and returns the independently computed correct value.
- Explanation: the answer states scope, assumptions, and limitations without inventing facts.
- Write intent: the right business method, preconditions, permission awareness, and required approval, avoiding dangerous raw writes.
Metrics
Primary metrics: exact executed-result correctness, semantic query-plan correctness, critical unsafe-action rate, and clarification correctness on ambiguous tasks. Secondary metrics include join and grain correctness, state/date/company/currency correctness, negative-knowledge adherence, hallucinated field or model rate, context token count, latency, and estimated cost, plus context precision and recall (were the assembled entries the ones actually needed).
Results are reported as absolute percentage-point change and raw counts, never as a lone relative percentage.
Ground truth and honesty controls
Ground truth is produced independently of the OCL entry wording: controlled fixtures with known expected values, Odoo business methods and reports where they are the authority, and independently reviewed queries, with result digests committed before the OCL entries are evaluated. For convention-dependent questions such as “revenue,” the expected correct behavior may be to ask a clarifying question, and scoring rewards that rather than forcing one universal definition.
Blind scoring is a review operation, not an LLM self-score. Answers are assigned opaque evaluation ids, the reviewer sees neither the case id nor which arm produced an answer, and the private map is restored only for deterministic scoring afterward.
Before any comparison is published
A claim is only published when the benchmark version and dates are stated, all arms used equivalent capabilities, task-family and aggregate results are shown with raw counts and run-to-run variance, failures and regressions are included, model ids and settings are disclosed, token and cost methodology is stated, and no case was added after seeing only the OCL result. The wording will say what was measured, never “AI now understands all of Odoo.”
When results clear that bar, they will appear here and on the Context Layer page. Until then, this page is the method and nothing more.