Models
One model for planning and review, one for execution. No vendor is named on this page, by policy.
AI control plane
Gevurah is an AI control plane with policy enforcement in code: multi-agent orchestration, persistent memory, and four gates that every outbound action must pass. It runs a real business in production.
The thesis
The market sells more capability, more agents, more autonomy. Gevurah was built around the other direction: gates that live in code, close by default, and write down why they closed. 29 of 137 policy decisions were blocked, and every block is stored with its reason.
Gevurah is the Hebrew word for the measure of restraint, the force that says no.
Policy decisions blocked102 days
0 / 137
Each blocked decision is stored with the reason it was blocked. The remaining 108 passed the gates and ran.
Category
Most tools in this space are libraries you assemble, or enterprise platforms you configure. Gevurah is neither. It is a system that has been running, on a schedule, against real money and real customers, and the interesting part is the moment an agent is told it does not get to run.
One model for planning and review, one for execution. No vendor is named on this page, by policy.
A large internal tool surface, built as the business needed it rather than as a product roadmap.
867 records, 1,315 wikilinks, 2,942 chunks across 751 files. Persistent across sessions and across agents.
1,195 work orders over 102 days, 929 results, 96 escalations to a human. 47 live coordination rooms.
4 gates, 42 read points, 39 write points, about 102 tests. A contact ledger with 281 records across 11 channels.
7,321 outbound actions logged over 92 days, and 484 uptime checks per day.
Autonomy
The scale runs L0 Observe, L1 Draft, L2 Prepare, L3 Bounded execute, L4 High-autonomy. The system runs at L3, and some flows at L4. An irreversible action automatically demotes the flow to L2, where a human approves before anything leaves the building. Money, publishing, and any message to a customer are irreversible by definition.
This is the answer to the question every operator actually asks. Capacity without the risk that the capacity acts on its own.
Evidence
1,195
delegated work orders over 102 days
Counted across the live queue and the archives.
44.5% to 63.4%
completed cleanly on the first attempt
The upper end counts 532 clean results out of the 839 tasks that have a linked result. The lower end counts the 356 work orders with no result record as failures, against all 1,195.
1.2%
full failure, 10 of the 839 measured
A further 21.1% completed partially, and 13.3% carry no outcome label at all.
96
escalations to a human over 102 days
An escalation is the system stopping and asking, which is the behaviour the gates exist to produce.
7,321
outbound actions logged over 92 days
A frozen snapshot taken on 12 August 2026. The ledger keeps writing.
29 / 137
policy decisions blocked, 21%
Each block is stored with the reason it was blocked.
The outcome field is written by the executing agent about its own run. This is self-assessment and not independent review. We publish the range rather than the upper number because the 356 work orders with no result record are a real gap, and treating them as successes would be a claim we cannot support.
A new runtime layer was built on 6 August 2026. The first six flows were migrated onto it from Windows Task Scheduler and ran 4,365 flows with 16 failures. The rest of the scheduled work is still on the old scheduler, and the migration is happening in waves.
A note on how that number nearly went out wrong. An earlier draft said the two pollers accounted for 84% of runs. 84% describes the two-minute poller on its own. Both pollers together are 99.1%. The measurement window is 5.24 days, and one flow inside it, the invoice pipeline, failed 5 times out of 8. The claim was rejected before it was published, by the same review the gates exist to force.
833 of the 839 measured tasks ran once only. 6 of 839 got a second attempt that already carried information from the first. pass^k cannot be derived from a historical record like that, so we built a harness that starts every attempt from a clean workspace, enforces a negative control, and requires an explicit reversibility approval. The harness is public under MIT at github.com/oreno334/passk. There is no canonical pass^k score for the system, and results will be reported separately.
The advertising the system runs returns close to a tenfold return as Google measures it. This is ad platform attribution and not audited sales.
Limits
Every claim on this page has a boundary, and the boundaries are written down in the same place the numbers are. What follows is the canonical limits record, published in full.
Tenant isolation. Tenant isolation is built and tested, and it has not been proven in production. 213 tests pass across 14 suites, and 11 of the 12 isolation guards are enforced. The twelfth guard waits on an explicit, documented decision. There is no paying external tenant.
Product and onboarding. This is not yet a self-serve product. The system is tailored to a single operator, and a new organisation would need specification and onboarding work before it could run this.
Reliability measurement. pass^k cannot be derived from the historical record, so we built a harness that starts every attempt from a clean workspace. 833 of the 839 measured tasks ran once only. Harness results will be reported separately.
Human oversight. The unsupervised dispatch layer has been off since 19 May. The operating mode is supervised concurrency, with a demotion to L2 before money, publishing, or a message to a customer.
Limits the audit exposed.
גבולות ההוכחה
ההסתייגויות כאן הן חלק מהראיה. הן מבהירות מה כבר נבדק, מה עדיין דורש הוכחת פרודקשן ואיפה אדם נשאר בתוך הלולאה.
בידוד דיירים בנוי ונבדק, אך טרם הוכח בפרודקשן. 213 טסטים עוברים ב-14 חבילות, ו-11 מתוך 12 הגנות הבידוד נאכפות. ההגנה הנוספת ממתינה להחלטה מפורשת ומתועדת. עדיין אין דייר חיצוני משלם.
זה עדיין אינו מוצר לרכישה עצמית. המערכת תפורה למפעיל אחד, וארגון חדש זקוק לעבודת אפיון והטמעה לפני שיוכל להפעיל אותה.
אי אפשר לגזור pass^k מהתיק ההיסטורי, ולכן בנינו הארנס שמתחיל כל ניסיון מסביבת עבודה נקייה. 833 מתוך 839 המשימות שנמדדו רצו פעם אחת בלבד. תוצאות ההארנס ידווחו בנפרד.
שכבת השיגור הלא-מפוקחת כבויה מאז 19 במאי. המצב התפעולי הוא מקביליות בפיקוח מפעיל, עם ירידה ל-L2 לפני כסף, פרסום או הודעה ללקוח.
Origin
OpticoAI is built on Gevurah's infrastructure and uses its technology. The agency work is both the funding and the proof: the founder funds the engineering from freelance and client work, and every gate in this system exists because something in that business was about to go wrong.
That is why the numbers on this page are operational rather than benchmarked. They come from a system that sends real messages, spends real budget, and occasionally gets stopped by its own policy layer.
Questions
A layer that sits above individual agents and decides what they are allowed to do. It holds orchestration, memory, and policy, so that capability and permission are separate concerns.
A framework is assembled per project and enforces nothing at runtime. A control plane is the thing that runs, keeps state between runs, and can refuse. The refusal is the part that is hard to retrofit.
L3 in production, and some flows at L4. Any irreversible action demotes the flow to L2 automatically, which puts a human in front of money, publishing, and customer messages.
Not yet. Tenant isolation is built and tested, with 213 tests passing across 14 suites and 11 of the 12 isolation guards enforced. The twelfth waits on an explicit documented decision. It has not been proven in production, and there is no paying external tenant.