Today we're opening Early Access to Konvu Guardrails.
Guardrails puts your security Controls inside the coding session. A model maps your Controls ahead of time, and code checks that every piece of evidence it cites exists in your repository.
When the coding agent is about to edit a file a Control covers, a small Rust binary looks up the matching Security Guidance: a short note on how this repository already handles that code, and where. The lookup runs locally, with no model and no network call.
That split between a slow, checked mapping step and a fast, deterministic edit path is most of the design. This post walks through it. First, the problem it's for.
Coding agents don't know your business rules
Coding agents now write a large share of the code, across many sessions at once, faster than anyone can review it. They know the textbook fixes. Ask one for a safe lookup and it will try to scope it.
What it doesn't know is your business rules: which tenant owns a record, which role may approve an action, which helper is your code's security control for a given table. Getting one of those wrong is a business-logic bug. As we wrote in September, these bugs look like valid code, so they often pass review, tests and scanners.
Scanners and pentesters find that kind of bug after it ships, and the next session writes it again. In one study of 1,186 vulnerabilities in 200 vibe-coded apps, researchers replayed the original feature request, and the agent wrote the same or a similar vulnerability again in about a quarter of runs. Finding them faster is still whack-a-mole.
Shift-left was supposed to fix this. It handed security work to busy developers with noisy tools, and they tuned out. Coding agents don't get alert fatigue, but they need your codebase's rules at the moment they write.
A rules file loaded at the start of a session doesn't deliver that. The agent may skim it, and it says nothing about the function it's about to change.
So we set four design goals:
- Specific to this codebase, with evidence from its own code.
- Delivered at the edit, and silent the rest of the time.
- No model on the edit path, so it's fast and predictable.
- Enforced before merge, for when steering isn't enough.
Guardrails: map your Controls, steer the agent, enforce at the PR
- Map. A model reads your code and proposes Controls, and the code that implements each one. Every claim has to point at real code, or it's dropped. The result is a Security Context Graph for the repository. Konvu can run the map for you, or you can run it locally.
- Steer. The first time the agent edits a file a Control covers, the hook refuses the edit once, with Security Guidance as the reason: what this repository requires for that code, the file and lines it's based on, and a request to say how the change follows it. The agent resubmits, changed or not, and the edit goes through.
- Enforce. A GitHub App runs a PR check. A model reviews the diff with the Security Context Graph as context. Make the check required, and a failing check blocks the merge.
- See. The dashboard shows your Controls, and which Controls steered which sessions on which endpoint. An endpoint is a developer machine running the coding agent, and each one appears by name.
Rolling it out is a managed setting in Claude Code. Nothing goes into your repositories.
Here is one example, then numbers showing that the failure it targets is common.
An illustrative run: one ticket, with and without guidance
We built this example to show the mechanism end to end. ihatemoney is an open-source app for sharing expenses. Its data is organized by project, and a project should never see another project's data. Several places in the app look people up through a project-scoped helper, Person.query.get(id, project).
We started from a copy of the app with that helper missing, and gave Claude Code the same ticket twice: add the project-scoped query methods to the Person model. The PR check was on both times. Security Guidance was on only for the second run.
Without Security Guidance, the agent wrote this:
def get(self, id, project=None):if not project:project = g.projectreturn (Person.query.filter(Person.id == id).filter(Project.id == project.id).one())
The project filter makes it read as scoped. But the query never joins Person to Project, so that condition only checks that the project exists, and the lookup returns the person with that ID whichever project they belong to. The Guardrails PR check failed with one finding: "Person lookups can cross project boundaries."
With Security Guidance, the agent's first edit of the model file came with guidance for the Control that covers it: every read of a Person in this app is scoped to the caller's project, with the file and line references it's based on. The agent wrote this:
def get(self, id, project=None):project = project or g.projecttry:return (self.filter(Person.id == id).filter(Person.project_id == project.id).one())except orm.exc.NoResultFound:return None
The filter is on the person's own project_id, so the lookup can only return someone from the caller's project. The test suite passed, and so did the PR check.
This is one run on one app. It shows how the mechanism works and says nothing about how often it works.
The failure itself is common. In September we ran Opus 5, with extended reasoning, on 184 feature-request tasks from the SusVibes benchmark, with no Guardrails in the loop. Each task comes from a real open-source project and is scored for both functionality and security. The model passed 96.2% of the functional checks and wrote vulnerable code on 72.3% of the tasks.
Code that works but isn't safe is the common case.
Think ahead, check instantly
Guardrails runs in three places, one for each of Map, Steer and Enforce. Map builds the Security Context Graph before any session needs it, either managed by Konvu or locally on your own machine. Steer runs a small hook on each developer's endpoint. Enforce runs the PR check on GitHub.
Map: a model proposes, your code proves
Every claim must point at real code
Business rules rarely look like security to a pattern matcher. A tenant check is often just a comparison that returns a bool, and its meaning lives in how the caller uses it. Recognizing that takes judgment, so mapping uses a model.
We don't trust its output directly. Mapping runs as a pipeline, and each stage checks the one before it:
- Index. Deterministic. It extracts facts: routes, functions, call sites and sensitive calls.
- Discover. The model reads those facts and names things: this function is a tenant check, this route reads a
Report, these call sites are the same Control written five different ways. A dedicated stage looks for business context, such as which records belong to which tenant and which roles gate which actions. - Compile. Deterministic again. Every piece of evidence the model cites has to resolve to locations the index knows about. Anything that doesn't is dropped.
The result is the Security Context Graph: the resources in the repository (an Invoice, a CompanyToken), the Controls that should hold for them, and the code that implements each Control.
The last step compiles the graph into a table of patterns, each paired with Security Guidance. That table is the only thing an endpoint receives. A pattern recognizes the files and code shapes that implement a Control, such as a query on the Person table or a particular config file. So the hook speaks when an edit touches one of those, not on every edit in the repository.
Where it runs: managed or local
Mapping is where a model reads your whole repository, so you choose where it runs. Both modes produce the same Security Context Graph, and the coding-agent plugin and the PR check use it the same way. (If you use the PR check, a model also reads each pull request in our cloud. More on that below.)
Managed. Konvu runs the map for you. The GitHub integration makes a shallow clone of the repository, and an OpenAI model deployment we operate proposes the Controls. If you'd rather the model calls run on your own account, bring your own OpenAI key.
Local. You run the map on your own machine with the Konvu CLI and your own key: konvu inventory map <path>. We never clone the repository, and the model calls go to OpenAI on your key.
Steer: a local hook in the coding agent
Rollout: one managed setting
A security team adds Guardrails once to Claude Code's managed settings, and every developer's next session picks it up.
What gets installed is a small, open-source plugin, so you can read exactly what runs on your machines. It wires up the hooks and downloads a pinned, checksum-verified guardrails binary that does the actual work.
Each machine enrolls itself with a short-lived deployment key and gets its own token. That token can only fetch your patterns and Security Guidance and report which guidance fired, so revoking a machine, or the whole rollout, is one action in Konvu.
Hooks: every edit checked locally
Guardrails works in three steps, each on a Claude Code hook:
| Step | Hook | When it fires | What Guardrails does |
|---|---|---|---|
| Fetch | SessionStart | A session opens or resumes | Fetches the Security Guidance in the background and keeps it locally. |
| Check | PreToolUse on Write, Edit, MultiEdit, NotebookEdit | Just before the agent writes a file | Checks the edit against the local Security Guidance. On a match, it refuses the first edit of that file once, with the guidance as the reason. The agent's retry goes through, and later edits get the guidance as advice. |
| Report | Stop | The agent finishes a turn | Uploads which guidance fired during the turn, without making the agent wait. |
No prompt data is sent to Konvu. The prompt is only read on the endpoint, to decide which Security Guidance applies.
Here are those hooks at work in the second session of the illustrative run, the one with Security Guidance. Scroll through it, or jump to a step.
When the session starts, Guardrails fetches your patterns and Security Guidance from Konvu, only if they changed, and keeps them on the machine. Nothing shows up in the session.
Before every write, a hook checks the edit against the local patterns. No model, no network call. No pattern matches, so the agent carries on undisturbed.
This is the first edit of the Person model. A pattern matches, so the hook refuses the edit once, with its Security Guidance as the reason: every read of a Person is scoped to the caller’s project, and the files and lines that show it.
The agent reads the guidance and resubmits the edit, filtering on the person’s own project_id. The edit goes through.
The agent opens the pull request. The tests pass, and the PR check reviews the diff with the Controls as context. The lookup is already scoped, so the check passes.
Reports: what fired, never prompts or code
Each report says which guidance was delivered, on which repository-relative file path, in which session and on which endpoint, and whether the edit was refused once or went through with the guidance as advice. The dashboard builds its view of which Controls steered which sessions from these reports.
Reports have no field for prompt text, the agent's responses or code, and our API rejects fields it doesn't expect.
Speed: tens of milliseconds
On the edit path, the hook does a local table lookup plus a few local git reads to find the repository. It answers in tens of milliseconds and calls no model, no code parser and no network.
Failure: fails open, but the agent can't turn it off
Guardrails sits in front of every edit, so it must never be the reason a developer is stuck. If it crashes, times out, can't reach Konvu, or has no guidance for the repository, the agent carries on as if it weren't installed.
You hold the off switch. Revoke an endpoint, or turn steering off for your company, and the hooks go quiet until the endpoint is authorized again. A developer who lost access sees one short "paused" notice, at most once a day.
The agent doesn't hold that switch. It can't edit Guardrails' rules or Claude Code's hook settings to turn Guardrails off.
Enforce: a PR check that can block the merge
Steering helps the agent get it right. The PR check confirms it did. A GitHub App reviews each pull request in our cloud.
The reviewer is a model, so we keep it grounded with the Security Context Graph: it knows which Controls apply to the diff and where your code already implements them.
The check passes or fails. Make it required in branch protection, and a failure blocks the merge. Leave it optional, and it only reports.
What's next: more agents, more sources, a benchmark
- More moments to steer. A hook when the agent writes its plan, not only when it edits, and coverage for files written through scripts.
- More coding agents. We're adding coding agents as Early Access users ask for them. Codex is first: it already runs Guardrails with manual setup, and next we package it the way we package Claude Code.
- Sources beyond code. Threat models, pentest reports and written policies.
- Repositories mapped together, so a Control that spans services is one Control.
- A published benchmark. We're measuring how much Security Guidance changes what coding agents ship. Early results are very encouraging, and we'll publish the full benchmark and method.
Where this goes: shift left that works
Shift left asked busy developers to act on noisy tools, at the wrong moment, and they tuned out. Coding agents change that. They don't get tired, they don't ignore alerts, and they follow guidance they can understand. What they lack is your context: the rules your code already follows, and where.
We believe that if you give agents that context just in time, at the edit, in a form they can follow, security stops being a step someone has to remember. It becomes part of how code gets written, and mostly invisible to the developer.
That's the bet behind Guardrails, and we're building it with our Early Access partners. Create an account, map a repository, and tell us where the guidance helps and where the map is wrong.