the independent audit for AI-generated code

Security audit for the code your AI assistant writes — here, or inside the assistant.

What you get →

Paste a repo or a .zip here, or let your assistant run the check over MCP. Either way: ship it, or fix these first. Every risk comes with evidence and a fix.

We publish every result — zero false alarms across 25 clean, mature libraries. Browse real scans in the report library, see what a scan reads, or set up your assistant in the docs.

Checking GitHub…

Free scan — no sign-in, no card — finds secrets and vulnerable packagesAudit — your first Audit is free. After that, about $5. Usually within 15 minutes*Report library See 86 free reports on real repos. Browse the library

The digest goes to Anthropic on nittim's key, kept up to 30 days. Bring your own key and it goes to your account.

* Most reports land within 15 minutes. Worst case, 24 hours.

See one real finding →

One finding, exactly as the report shows it

GitHub Actions script injection via release event fields in news-update workflows

HighSecuritySmall (hours)
Evidence
.github/workflows/update-news-www.yml:19and .github/workflows/update-news-www-legacy.yml:19 embed ${{ github.event.release.published_at }} and ${{ github.event.release.tag_name }} directly into a shell `sed -i` command (IaC scan: gha-script-injection).
Impact
Untrusted-looking event input interpolated into a shell command can allow command injection into the CI runner, potentially compromising build secrets or the project's own web content pipeline — a supply-chain integrity risk distinct from the app's intentional vulnerabilities.
Cause
Direct interpolation of GitHub event context into a run: shell step instead of passing values via environment variables.
Fix
Assign the event fields to env: variables and reference them as "$TAG_NAME"/"$PUBLISHED_AT" inside the script, or validate/escape them; never inline ${{ github.event.* }} into shell.
RemediationSmall (hours)

From the report on juice-shop/juice-shop. Read the whole report →

86 open-source reports, free. Browse the library

How it works

01

Point it at a repo

A public URL, a repo, or a zip. Free scans need no account. Files leave your machine and are read in nittim's server memory. No model, no verdict. Kept: your report, with the line each finding points at — off by setting.

02

We read what we can, honestly

Coverage is stated up front, in the report — not a footnote.

03

Get scored, sourced findings

Severity, evidence, a fix for each — or Full Audit reads everything. The digest goes to Anthropic on nittim's key, kept up to 30 days. Bring your own key and it goes to your account.

Why a green report means something

The author never grades the paper. An independent audit scores your code against one fixed framework.

A floor the model can't sink below. Deterministic scanners catch committed secrets and known CVEs, not moods.

Fix, recheck, repeat — no babysitting. Findings go to your assistant, it fixes them, the audit reruns. Review a closed loop, not a to-do list.

Every claim comes with evidence. File and line. Dispute one, get the quote back — not “trust me.”

Production-safe is a state, not a feeling. 13 categories, six scores, one verdict — ship or don't, with reasons.

See the full framework →

The Loop — free, no code sent

A checklist your own assistant runs, where your code already is.

Nothing of your code is sent to nittim. Your assistant runs it where it already works. Anonymised counts only, and only if you choose.

In the machine’s own words

Read it:

Let your AI assistant do it

Free, forever. Ask your assistant — no key, no account. Checks secrets and known CVEs; save to CLAUDE.md and teammates get it too.

Paste this to your assistant:

I’d like to connect nittim’s code scanner to this project. nittim is a code-audit service at nittim.com; what it does with code it reads is at nittim.com/trust.

So you don’t have to take a vendor’s word for it: the scan works by sending file contents to nittim.com. I’ve read their trust page and decided that’s acceptable for this repo — that call is mine, not yours.

It runs as an MCP server, so this adds one entry to my client’s config — not to my repo — and my client still asks me before a tool call. With no key it can read and price: scan_source, list_modules and estimate_audit, and none of it costs anything.

Please connect it:
claude mcp add --transport http nittim https://nittim.com/api/mcp

Then, before you scan this folder for committed secrets and known vulnerable dependencies, show me the list of files you would send and wait for my yes. Leave out anything holding customer data. I have no nittim key, so skip anything that costs money.

Free — no account, no nittim credits. Same deterministic scan as the free web scan: no AI, read in memory, keeps findings, never files. Commit it to CLAUDE.md and every teammate's assistant picks it up automatically — no key in the file, because none exists to leak. Get the CLAUDE.md block →

Want the deeper, 13-category self-review — on your own model, nothing sent to nittim? Run it as the MCP prompt nittim-selfcheck once connected, or directly: /selfcheck

What we check

14 things that decide whether your software survives real users — not just syntax.

Linters check syntax. nittim reasons about whether your software survives production. The module set is chosen per repo from its file tree; every report's “Modules run” strip names exactly what fired.

Six scores — Executive, Production Readiness, Security, Privacy, Architecture, IP Protection — and one verdict: Production Ready · Production Ready with Conditions · High Risk · Not Safe for Production. The verdict rests on hard evidence in your code — confident writing can't lift it.

Can someone break in?

Security · counts as Core toward the score

We look for the vulnerabilities that turn a data breach into a headline — injection flaws, broken authentication, exposed secrets, and gaps in how your APIs and infrastructure protect themselves. Every finding points to the exact file and line, with a fix.

module security · priority tier Core · Claude-reasoned · findings land under Security

Are you handling people's data properly?

Privacy & Compliance · counts as Core toward the score

We check how your app collects, stores, and hands off personal data — consent, data governance, audit trails, and readiness for the regulations that actually apply to you, including GDPR, CCPA, HIPAA, PCI DSS, and SOC 2.

module privacy · priority tier Core · Claude-reasoned · findings land under Privacy & Compliance

What happens when something goes wrong?

Reliability & Resilience · counts as Core toward the score

We look at what happens when something goes wrong — a crash, a dropped connection, a bad deploy — and whether your app recovers gracefully or takes your users down with it. Backups, failover, and fault tolerance all fall here.

module reliability · priority tier Core · Claude-reasoned · findings land under Reliability & Resilience

Can you build on this next month?

Code Quality & Architecture · counts as Important toward the score

We read your codebase the way a senior engineer would on day one: is it organized, is it readable, and how much will the next feature cost to build on top of it. This is where technical debt gets named.

module code_quality · priority tier Important · Claude-reasoned · findings land under Code Quality & Architecture

Did the AI fake anything?

AI / Vibe Coding Risk · counts as Important toward the score

AI-generated code has its own failure modes — confident-looking placeholder logic, hallucinated APIs, copy-pasted duplication, and edge cases nobody thought to handle because nobody wrote the code by hand. We look specifically for these.

module ai_risk · priority tier Important · Claude-reasoned · findings land under AI / Vibe Coding Risk

Will it slow down when people show up?

Performance & Scalability · counts as Important toward the score

We look at what slows your app down under real load — inefficient queries, missing caching, wasted memory, and anything that gets more expensive as you grow.

module performance · priority tier Important · Claude-reasoned · findings land under Performance & Scalability

Can you deploy it safely, again and again?

Infrastructure & DevOps · counts as Important toward the score

We check the machinery that gets your code into production and keeps it there — your CI/CD pipeline, container and infrastructure setup, environment separation, and how secrets are managed through their whole lifecycle.

module devops · priority tier Important · Claude-reasoned · findings land under Infrastructure & DevOps

Could your data get quietly corrupted?

Data Layer · counts as Important toward the score

We look at your data layer for the mistakes that are hardest to undo — unsafe migrations, missing constraints, unencrypted sensitive fields, and transaction handling that can silently corrupt data.

module data · priority tier Important · Claude-reasoned · findings land under Data Layer

Can you run a business on it?

Business & Product Risk · counts as Supporting toward the score

We step back and ask the practical questions a founder or investor would: is this ready to run a real business on, what does it actually cost to keep alive, and where are you locked into a vendor or a license you didn't mean to take on.

module business · priority tier Supporting · Claude-reasoned · findings land under Business & Product Risk

Is it painful to work in?

Developer Experience · counts as Supporting toward the score

We check how easy your own repository is to work in — local setup, build reproducibility, onboarding a new engineer, and the release process. A repo that's painful to develop in slows down everything else.

module devex · priority tier Supporting · Claude-reasoned · findings land under Developer Experience

Can everyone use it?

Accessibility & UX · counts as Supporting toward the score

We check whether your app is usable by everyone — screen-reader and keyboard support, responsive layouts, clear error messaging, and basic internationalization and browser-compatibility gaps.

module accessibility · priority tier Supporting · Claude-reasoned · findings land under Accessibility & UX

Would you notice if it broke?

Observability · counts as Supporting toward the score

We check whether you'd actually notice if something broke — logging, metrics, tracing, and alerting that gets a human to the right place before your users notice first.

module observability · priority tier Supporting · Claude-reasoned · findings land under Observability

Will it still be maintainable in a year?

Maintainability Forecast · counts as Specialized toward the score

We take a forward-looking view: how much of this codebase depends on one person's memory, how expensive a future rewrite would be, and whether this is something you can keep maintaining a year from now.

module maintainability · priority tier Specialized · Claude-reasoned · findings land under Maintainability Forecast

Are your best ideas showing?

IP & Novelty Exposure · counts as Specialized toward the score

We flag the genuinely novel ideas in your codebase — algorithms, architectures, or methods worth protecting — and check two separate ways they can leak: through a public repository, and through public-facing copy (README, docs, marketing) that spells out how something works rather than just what it does. This is an independent signal and never affects your production verdict.

module ip_exposure · priority tier Specialized · Claude-reasoned · findings land under IP & Novelty Exposure

And if you handle personal data: GDPR readiness

Runs automatically when a repository shows signs of handling personal data.

A focused pass on GDPR readiness: whether you have a lawful basis for the personal data you collect, real consent flows, a clear picture of where personal data lives, a working path to honor deletion and data-subject-rights requests, and a plan for cross-border transfers and breach notification. Findings here are grouped with the broader Privacy & Compliance check. nittim runs this automatically when a repository shows signs of handling personal data.

How often we’re wrong

We measure it and publish it — with the confidence interval and the caveats attached. 0 false alarms on 25 clean libraries (95% CI 0–13.8%).
Measured, not claimed
0 of 25
false criticals — clean, mature libraries, both tiers
11 of 11
intentionally-vulnerable apps verdicted unsafe
0 of 20
secret false-criticals — real, deployed apps
16 of 20
mature deployed apps passed green, Deep tier

Small corpora, honest caveats — every number ships with its method and confidence interval.

0 of 25 false criticals → 95% CI 0–13.8% (Clopper-Pearson). Verdict-level recall 100% (11 of 11); 91.7% under strict pre-registration. The libraries corpus is IN-SAMPLE — four precision rules were written against its failures — so read it beside the 12-library held-out number (1 of 12 secret FCR, pre-registered before it ran), never alone. The judge shares the Opus family; only the separate MCP cross-check is cross-vendor. Every not-green deployed-app verdict was adjudicated over-escalated on dependency drift; root causes fixed, re-run outstanding.

A human audit runs $500–$3,000 and takes days to weeks. nittim deep-verifies a repo in minutes — evidence included.

Is my code safe with you?

Start without showing a single line of code. Share more only when you decide to.

You shouldn't have to trust nittim before you get anything from it. Start at the top.

0 · Your assistant, our rubric

Nothing reaches nittim. No independent judge either.

Run /selfcheck’s 13 categories with your own assistant — you give us nothing. It wrote the code, so it may repeat its own blind spots; an independent judge catches those, and that’s what we sell.

1 · Our report library

Nothing of yours leaves your machine.

Read real scans of well-known public code in /library before nittim sees a line of yours. Nothing to take on faith.

2 · Your repo, free scan

Read in memory, never written down.

No AI in the free scan. Two deterministic scanners read the files and build the report. A test fails the build if source text reaches the database — a property of the code, not a promise.

3 · AI audit on your own key

The reasoning runs on your Anthropic account.

BYOK Pro — API, MCP, or GitHub Action: the call runs on your key, under your vendor terms. nittim's model account never sees your code — check your own Anthropic bill.

4 · AI audit, on our key

Your code, in memory, plus our model vendor.

Held in memory for the audit, never trained on, discarded after. The one rung that asks you to take nittim's word for it — last on purpose.

5 · Your own infrastructure

Nothing leaves your cloud at all.

Built, both halves ran end to end, and it runs in your account — no source, finding or repo name reaches nittim. Not self-serve: we set it up with you on request. Images aren’t signed yet — verify by published hash.

Stored
Audit report — findings, scores, verdict, repository name.
Not stored
Source code — fetched into server memory during analysis, discarded when the pipeline finishes. Exception: finding evidence includes the lines it points to.
Not stored
GitHub tokens — used once, discarded. Installation tokens expire within the hour; nittim keeps only the installation id.
Read-only
GitHub App: exactly two permissions — contents and metadata, both read-only — no events. nittim can never modify your code.

What it costs

Scanning is free, forever. A GitHub or Google account gets one covered Audit. After that, $29 buys two. The digest goes to Anthropic on nittim's key, kept up to 30 days. Bring your own key and it goes to your account.

nittim never bills flat rate for tokens it doesn't control — prepaid credits, your own Anthropic key, or your own cloud, your choice, never a surprise.

Free
Deterministic secret + dependency-CVE scan — any public or connected repo, forever. No card, no sign-in. One covered Audit run per account (GitHub or Google).
Credits
1 credit = $1. An AI audit costs 5.14 credits; a single module via API/MCP costs 5.03. Packs: $29 → 5 · $74 → 17 · $189 → 52 · $479 → 157 audits. Credits never expire · unused credits refundable within 30 days.
BYOK Pro
$49/mo pays for the pipeline, not tokens — bring your own Anthropic key. A 100-credit monthly allowance covers ~50 single-pass audits at 2 credits each.
Enterprise
from $18k/yr — seat bands, unlimited repos, SSO, SARIF, policy-as-code, append-only audit log. Runs on Claude via Bedrock in your AWS, or your own Anthropic org. Orgs share one audit portfolio and one bill. Talk to us →