Steer in the session, enforce at the PR: how Konvu Guardrails works

    Lucas Masson
    2026-10-08

    Today we're opening Early Access to Konvu Guardrails.

    Guardrails puts your security Controls inside the coding session. A model maps your Controls ahead of time, and code checks that every piece of evidence it cites exists in your repository.

    When the coding agent is about to edit a file a Control covers, a small Rust binary looks up the matching Security Guidance: a short note on how this repository already handles that code, and where. The lookup runs locally, with no model and no network call.

    That split between a slow, checked mapping step and a fast, deterministic edit path is most of the design. This post walks through it. First, the problem it's for.

    Coding agents don't know your business rules

    Coding agents now write a large share of the code, across many sessions at once, faster than anyone can review it. They know the textbook fixes. Ask one for a safe lookup and it will try to scope it.

    What it doesn't know is your business rules: which tenant owns a record, which role may approve an action, which helper is your code's security control for a given table. Getting one of those wrong is a business-logic bug. As we wrote in September, these bugs look like valid code, so they often pass review, tests and scanners.

    Scanners and pentesters find that kind of bug after it ships, and the next session writes it again. In one study of 1,186 vulnerabilities in 200 vibe-coded apps, researchers replayed the original feature request, and the agent wrote the same or a similar vulnerability again in about a quarter of runs. Finding them faster is still whack-a-mole.

    Fixing business-logic bugs after they ship: knock one down, and the next coding session puts another one up.

    Shift-left was supposed to fix this. It handed security work to busy developers with noisy tools, and they tuned out. Coding agents don't get alert fatigue, but they need your codebase's rules at the moment they write.

    A rules file loaded at the start of a session doesn't deliver that. The agent may skim it, and it says nothing about the function it's about to change.

    So we set four design goals:

    • Specific to this codebase, with evidence from its own code.
    • Delivered at the edit, and silent the rest of the time.
    • No model on the edit path, so it's fast and predictable.
    • Enforced before merge, for when steering isn't enough.

    Guardrails: map your Controls, steer the agent, enforce at the PR

    Your security context, across the AI SDLC. Konvu's Security Context Graph holds your Controls, mapped from your code. In a coding agent session on a developer workstation or background agent: 1, at SessionStart the hook fetches patterns and Security Guidance; 2, they are synced to the endpoint only when they change; 3, before each write the edit is matched locally, with guidance on a match. 4, in CI/CD, the pull request gets a PR check with your Controls, before merge and deploy.
    Konvu builds the Security Context Graph. Coding sessions and pull requests use it.
    1. Map. A model reads your code and proposes Controls, and the code that implements each one. Every claim has to point at real code, or it's dropped. The result is a Security Context Graph for the repository. Konvu can run the map for you, or you can run it locally.
    2. Steer. The first time the agent edits a file a Control covers, the hook refuses the edit once, with Security Guidance as the reason: what this repository requires for that code, the file and lines it's based on, and a request to say how the change follows it. The agent resubmits, changed or not, and the edit goes through.
    3. Enforce. A GitHub App runs a PR check. A model reviews the diff with the Security Context Graph as context. Make the check required, and a failing check blocks the merge.
    4. See. The dashboard shows your Controls, and which Controls steered which sessions on which endpoint. An endpoint is a developer machine running the coding agent, and each one appears by name.
    The Guardrails dashboard. A mapped repository has 24 Controls learned, grouped into control packs such as Validation and Business Logic and Authorization. Steering for coding agents and the PR check are both switched on, and each control pack shows where it is implemented, how often it steered agents, and how many issues the PR check caught and got fixed.
    The dashboard: Controls learned from your code, and where they steer and enforce.

    Rolling it out is a managed setting in Claude Code. Nothing goes into your repositories.

    Here is one example, then numbers showing that the failure it targets is common.

    An illustrative run: one ticket, with and without guidance

    We built this example to show the mechanism end to end. ihatemoney is an open-source app for sharing expenses. Its data is organized by project, and a project should never see another project's data. Several places in the app look people up through a project-scoped helper, Person.query.get(id, project).

    We started from a copy of the app with that helper missing, and gave Claude Code the same ticket twice: add the project-scoped query methods to the Person model. The PR check was on both times. Security Guidance was on only for the second run.

    Without Security Guidance, the agent wrote this:

    def get(self, id, project=None):
    if not project:
    project = g.project
    return (
    Person.query.filter(Person.id == id)
    .filter(Project.id == project.id)
    .one()
    )

    The project filter makes it read as scoped. But the query never joins Person to Project, so that condition only checks that the project exists, and the lookup returns the person with that ID whichever project they belong to. The Guardrails PR check failed with one finding: "Person lookups can cross project boundaries."

    With Security Guidance, the agent's first edit of the model file came with guidance for the Control that covers it: every read of a Person in this app is scoped to the caller's project, with the file and line references it's based on. The agent wrote this:

    def get(self, id, project=None):
    project = project or g.project
    try:
    return (
    self.filter(Person.id == id)
    .filter(Person.project_id == project.id)
    .one()
    )
    except orm.exc.NoResultFound:
    return None

    The filter is on the person's own project_id, so the lookup can only return someone from the caller's project. The test suite passed, and so did the PR check.

    This is one run on one app. It shows how the mechanism works and says nothing about how often it works.

    The failure itself is common. In September we ran Opus 5, with extended reasoning, on 184 feature-request tasks from the SusVibes benchmark, with no Guardrails in the loop. Each task comes from a real open-source project and is scored for both functionality and security. The model passed 96.2% of the functional checks and wrote vulnerable code on 72.3% of the tasks.

    Code that works isn't code that's safe. On 184 SusVibes feature requests from real open-source projects, Opus 5 with extended reasoning passed 96% of functional checks (177 of 184) but only 28% of security checks (51 of 184).
    Opus 5 on 184 SusVibes tasks, run by Konvu in September 2026.

    Code that works but isn't safe is the common case.

    Think ahead, check instantly

    Guardrails runs in three places, one for each of Map, Steer and Enforce. Map builds the Security Context Graph before any session needs it, either managed by Konvu or locally on your own machine. Steer runs a small hook on each developer's endpoint. Enforce runs the PR check on GitHub.

    Map: a model proposes, your code proves

    Every claim must point at real code

    Business rules rarely look like security to a pattern matcher. A tenant check is often just a comparison that returns a bool, and its meaning lives in how the caller uses it. Recognizing that takes judgment, so mapping uses a model.

    We don't trust its output directly. Mapping runs as a pipeline, and each stage checks the one before it:

    1. Index. Deterministic. It extracts facts: routes, functions, call sites and sensitive calls.
    2. Discover. The model reads those facts and names things: this function is a tenant check, this route reads a Report, these call sites are the same Control written five different ways. A dedicated stage looks for business context, such as which records belong to which tenant and which roles gate which actions.
    3. Compile. Deterministic again. Every piece of evidence the model cites has to resolve to locations the index knows about. Anything that doesn't is dropped.

    The result is the Security Context Graph: the resources in the repository (an Invoice, a CompanyToken), the Controls that should hold for them, and the code that implements each Control.

    The last step compiles the graph into a table of patterns, each paired with Security Guidance. That table is the only thing an endpoint receives. A pattern recognizes the files and code shapes that implement a Control, such as a query on the Person table or a particular config file. So the hook speaks when an edit touches one of those, not on every edit in the repository.

    Where it runs: managed or local

    Mapping is where a model reads your whole repository, so you choose where it runs. Both modes produce the same Security Context Graph, and the coding-agent plugin and the PR check use it the same way. (If you use the PR check, a model also reads each pull request in our cloud. More on that below.)

    Managed. Konvu runs the map for you. The GitHub integration makes a shallow clone of the repository, and an OpenAI model deployment we operate proposes the Controls. If you'd rather the model calls run on your own account, bring your own OpenAI key.

    Choosing repositories to map from the connected GitHub integration. Each repository shows its asset tier and threat signals, key assets are listed first, and ihatemoney is already mapped.
    Choosing which repositories to map, from the GitHub integration.

    Local. You run the map on your own machine with the Konvu CLI and your own key: konvu inventory map <path>. We never clone the repository, and the model calls go to OpenAI on your key.

    Steer: a local hook in the coding agent

    Rollout: one managed setting

    A security team adds Guardrails once to Claude Code's managed settings, and every developer's next session picks it up.

    What gets installed is a small, open-source plugin, so you can read exactly what runs on your machines. It wires up the hooks and downloads a pinned, checksum-verified guardrails binary that does the actual work.

    Each machine enrolls itself with a short-lived deployment key and gets its own token. That token can only fetch your patterns and Security Guidance and report which guidance fired, so revoking a machine, or the whole rollout, is one action in Konvu.

    Hooks: every edit checked locally

    Guardrails works in three steps, each on a Claude Code hook:

    StepHookWhen it firesWhat Guardrails does
    FetchSessionStartA session opens or resumesFetches the Security Guidance in the background and keeps it locally.
    CheckPreToolUse on Write, Edit, MultiEdit, NotebookEditJust before the agent writes a fileChecks the edit against the local Security Guidance. On a match, it refuses the first edit of that file once, with the guidance as the reason. The agent's retry goes through, and later edits get the guidance as advice.
    ReportStopThe agent finishes a turnUploads which guidance fired during the turn, without making the agent wait.

    No prompt data is sent to Konvu. The prompt is only read on the endpoint, to decide which Security Guidance applies.

    Here are those hooks at work in the second session of the illustrative run, the one with Security Guidance. Scroll through it, or jump to a step.

    coding agent · ~/ihatemoney
    > Add the project-scoped query methods to the Person model
    ⏺ Update(CHANGELOG.rst)
    ⎿ Added 1 line
    ⏺ Update(ihatemoney/models.py)
    + def get(self, id, project=None):
    ⎿ Refused once · Konvu Guardrails
    Every read of a Person is scoped to the caller’s project. Filter on the person’s own project, like the other project-scoped lookups in this app.
    ⏺ Update(ihatemoney/models.py)
    + .filter(Person.project_id == project.id)
    ⏺ Done. Person lookups are scoped to the project.
    ⏺ Bash(gh pr create --fill)
    ⎿ Tests: passed · PR check · Konvu Guardrails: passed

    When the session starts, Guardrails fetches your patterns and Security Guidance from Konvu, only if they changed, and keeps them on the machine. Nothing shows up in the session.

    Konvu→Patterns + guidance, on the endpoint

    Before every write, a hook checks the edit against the local patterns. No model, no network call. No pattern matches, so the agent carries on undisturbed.

    Write→Hook→Local patterns→nothing to say

    This is the first edit of the Person model. A pattern matches, so the hook refuses the edit once, with its Security Guidance as the reason: every read of a Person is scoped to the caller’s project, and the files and lines that show it.

    Update→Hook→Local patterns→Security Guidance

    The agent reads the guidance and resubmits the edit, filtering on the person’s own project_id. The edit goes through.

    Security Guidance→Scoped edit

    The agent opens the pull request. The tests pass, and the PR check reviews the diff with the Controls as context. The lookup is already scoped, so the check passes.

    Pull request→PR check→passes

    Reports: what fired, never prompts or code

    Each report says which guidance was delivered, on which repository-relative file path, in which session and on which endpoint, and whether the edit was refused once or went through with the guidance as advice. The dashboard builds its view of which Controls steered which sessions from these reports.

    One Control in the dashboard: "Project record lookups and collections are constrained to the current project." Over the last 30 days it steered coding agents 8 times, and the PR check checked 3 pull requests, all fixed after the check and none merged anyway. Recent activity lists those pull requests and a coding-agent session it steered on a named endpoint, and the definition lists the code locations that implement the Control.
    One Control: how often it steered coding agents, and what the PR check caught.

    Reports have no field for prompt text, the agent's responses or code, and our API rejects fields it doesn't expect.

    Speed: tens of milliseconds

    On the edit path, the hook does a local table lookup plus a few local git reads to find the repository. It answers in tens of milliseconds and calls no model, no code parser and no network.

    Failure: fails open, but the agent can't turn it off

    Guardrails sits in front of every edit, so it must never be the reason a developer is stuck. If it crashes, times out, can't reach Konvu, or has no guidance for the repository, the agent carries on as if it weren't installed.

    You hold the off switch. Revoke an endpoint, or turn steering off for your company, and the hooks go quiet until the endpoint is authorized again. A developer who lost access sees one short "paused" notice, at most once a day.

    The agent doesn't hold that switch. It can't edit Guardrails' rules or Claude Code's hook settings to turn Guardrails off.

    Enforce: a PR check that can block the merge

    Steering helps the agent get it right. The PR check confirms it did. A GitHub App reviews each pull request in our cloud.

    The reviewer is a model, so we keep it grounded with the Security Context Graph: it knows which Controls apply to the diff and where your code already implements them.

    The check passes or fails. Make it required in branch protection, and a failure blocks the merge. Leave it optional, and it only reports.

    The Konvu PR check commenting on a pull request: check failed, with one Authorization finding, "Person lookups can cross project boundaries", flagged as an introduced security regression. The explanation says the new PersonQuery.get lookup doesn't relate the requested project to the person, so a user of one project can edit, delete or deactivate a person in another project by supplying that person's ID.
    The PR check on the illustrative run without Security Guidance: one finding, and the check fails.

    What's next: more agents, more sources, a benchmark

    • More moments to steer. A hook when the agent writes its plan, not only when it edits, and coverage for files written through scripts.
    • More coding agents. We're adding coding agents as Early Access users ask for them. Codex is first: it already runs Guardrails with manual setup, and next we package it the way we package Claude Code.
    • Sources beyond code. Threat models, pentest reports and written policies.
    • Repositories mapped together, so a Control that spans services is one Control.
    • A published benchmark. We're measuring how much Security Guidance changes what coding agents ship. Early results are very encouraging, and we'll publish the full benchmark and method.

    Where this goes: shift left that works

    Shift left asked busy developers to act on noisy tools, at the wrong moment, and they tuned out. Coding agents change that. They don't get tired, they don't ignore alerts, and they follow guidance they can understand. What they lack is your context: the rules your code already follows, and where.

    We believe that if you give agents that context just in time, at the edit, in a form they can follow, security stops being a step someone has to remember. It becomes part of how code gets written, and mostly invisible to the developer.

    That's the bet behind Guardrails, and we're building it with our Early Access partners. Create an account, map a repository, and tell us where the guidance helps and where the map is wrong.