the independent audit for AI-generated code
Security audit for the code your AI assistant writes — here, or inside the assistant.
What you get →
Paste a repo or a .zip here, or let your assistant run the check over MCP. Either way: ship it, or fix these first. Every risk comes with evidence and a fix.
We publish every result — zero false alarms across 25 clean, mature libraries. Browse real scans in the report library, see what a scan reads, or set up your assistant in the docs.
See one real finding →
One finding, exactly as the report shows it
GitHub Actions script injection via release event fields in news-update workflows
- Evidence
.github/workflows/update-news-www.yml:19and .github/workflows/update-news-www-legacy.yml:19 embed ${{ github.event.release.published_at }} and ${{ github.event.release.tag_name }} directly into a shell `sed -i` command (IaC scan: gha-script-injection).- Impact
- Untrusted-looking event input interpolated into a shell command can allow command injection into the CI runner, potentially compromising build secrets or the project's own web content pipeline — a supply-chain integrity risk distinct from the app's intentional vulnerabilities.
- Cause
- Direct interpolation of GitHub event context into a run: shell step instead of passing values via environment variables.
- Fix
- Assign the event fields to env: variables and reference them as "$TAG_NAME"/"$PUBLISHED_AT" inside the script, or validate/escape them; never inline ${{ github.event.* }} into shell.
From the report on juice-shop/juice-shop. Read the whole report →
86 open-source reports, free. Browse the library
How it works
Point it at a repo
A public URL, a repo, or a zip. Free scans need no account. Files leave your machine and are read in nittim's server memory. No model, no verdict. Kept: your report, with the line each finding points at — off by setting.
We read what we can, honestly
Coverage is stated up front, in the report — not a footnote.
Get scored, sourced findings
Severity, evidence, a fix for each — or Full Audit reads everything. The digest goes to Anthropic on nittim's key, kept up to 30 days. Bring your own key and it goes to your account.
Why a green report means something
The author never grades the paper. An independent audit scores your code against one fixed framework.
A floor the model can't sink below. Deterministic scanners catch committed secrets and known CVEs, not moods.
Fix, recheck, repeat — no babysitting. Findings go to your assistant, it fixes them, the audit reruns. Review a closed loop, not a to-do list.
Every claim comes with evidence. File and line. Dispute one, get the quote back — not “trust me.”
Production-safe is a state, not a feeling. 13 categories, six scores, one verdict — ship or don't, with reasons.
The Loop — free, no code sent
A checklist your own assistant runs, where your code already is.
In the machine’s own words
Let your AI assistant do it
Paste this to your assistant:
I’d like to connect nittim’s code scanner to this project. nittim is a code-audit service at nittim.com; what it does with code it reads is at nittim.com/trust. So you don’t have to take a vendor’s word for it: the scan works by sending file contents to nittim.com. I’ve read their trust page and decided that’s acceptable for this repo — that call is mine, not yours. It runs as an MCP server, so this adds one entry to my client’s config — not to my repo — and my client still asks me before a tool call. With no key it can read and price: scan_source, list_modules and estimate_audit, and none of it costs anything. Please connect it: claude mcp add --transport http nittim https://nittim.com/api/mcp Then, before you scan this folder for committed secrets and known vulnerable dependencies, show me the list of files you would send and wait for my yes. Leave out anything holding customer data. I have no nittim key, so skip anything that costs money.
Free — no account, no nittim credits. Same deterministic scan as the free web scan: no AI, read in memory, keeps findings, never files. Commit it to CLAUDE.md and every teammate's assistant picks it up automatically — no key in the file, because none exists to leak. Get the CLAUDE.md block →
Want the deeper, 13-category self-review — on your own model, nothing sent to nittim? Run it as the MCP prompt nittim-selfcheck once connected, or directly: /selfcheck
What we check
Linters check syntax. nittim reasons about whether your software survives production. The module set is chosen per repo from its file tree; every report's “Modules run” strip names exactly what fired.
Six scores — Executive, Production Readiness, Security, Privacy, Architecture, IP Protection — and one verdict: Production Ready · Production Ready with Conditions · High Risk · Not Safe for Production. The verdict rests on hard evidence in your code — confident writing can't lift it.
Can someone break in?
We look for the vulnerabilities that turn a data breach into a headline — injection flaws, broken authentication, exposed secrets, and gaps in how your APIs and infrastructure protect themselves. Every finding points to the exact file and line, with a fix.
module security · priority tier Core · Claude-reasoned · findings land under Security
Are you handling people's data properly?
We check how your app collects, stores, and hands off personal data — consent, data governance, audit trails, and readiness for the regulations that actually apply to you, including GDPR, CCPA, HIPAA, PCI DSS, and SOC 2.
module privacy · priority tier Core · Claude-reasoned · findings land under Privacy & Compliance
What happens when something goes wrong?
We look at what happens when something goes wrong — a crash, a dropped connection, a bad deploy — and whether your app recovers gracefully or takes your users down with it. Backups, failover, and fault tolerance all fall here.
module reliability · priority tier Core · Claude-reasoned · findings land under Reliability & Resilience
Can you build on this next month?
We read your codebase the way a senior engineer would on day one: is it organized, is it readable, and how much will the next feature cost to build on top of it. This is where technical debt gets named.
module code_quality · priority tier Important · Claude-reasoned · findings land under Code Quality & Architecture
Did the AI fake anything?
AI-generated code has its own failure modes — confident-looking placeholder logic, hallucinated APIs, copy-pasted duplication, and edge cases nobody thought to handle because nobody wrote the code by hand. We look specifically for these.
module ai_risk · priority tier Important · Claude-reasoned · findings land under AI / Vibe Coding Risk
Will it slow down when people show up?
We look at what slows your app down under real load — inefficient queries, missing caching, wasted memory, and anything that gets more expensive as you grow.
module performance · priority tier Important · Claude-reasoned · findings land under Performance & Scalability
Can you deploy it safely, again and again?
We check the machinery that gets your code into production and keeps it there — your CI/CD pipeline, container and infrastructure setup, environment separation, and how secrets are managed through their whole lifecycle.
module devops · priority tier Important · Claude-reasoned · findings land under Infrastructure & DevOps
Could your data get quietly corrupted?
We look at your data layer for the mistakes that are hardest to undo — unsafe migrations, missing constraints, unencrypted sensitive fields, and transaction handling that can silently corrupt data.
module data · priority tier Important · Claude-reasoned · findings land under Data Layer
Can you run a business on it?
We step back and ask the practical questions a founder or investor would: is this ready to run a real business on, what does it actually cost to keep alive, and where are you locked into a vendor or a license you didn't mean to take on.
module business · priority tier Supporting · Claude-reasoned · findings land under Business & Product Risk
Is it painful to work in?
We check how easy your own repository is to work in — local setup, build reproducibility, onboarding a new engineer, and the release process. A repo that's painful to develop in slows down everything else.
module devex · priority tier Supporting · Claude-reasoned · findings land under Developer Experience
Can everyone use it?
We check whether your app is usable by everyone — screen-reader and keyboard support, responsive layouts, clear error messaging, and basic internationalization and browser-compatibility gaps.
module accessibility · priority tier Supporting · Claude-reasoned · findings land under Accessibility & UX
Would you notice if it broke?
We check whether you'd actually notice if something broke — logging, metrics, tracing, and alerting that gets a human to the right place before your users notice first.
module observability · priority tier Supporting · Claude-reasoned · findings land under Observability
Will it still be maintainable in a year?
We take a forward-looking view: how much of this codebase depends on one person's memory, how expensive a future rewrite would be, and whether this is something you can keep maintaining a year from now.
module maintainability · priority tier Specialized · Claude-reasoned · findings land under Maintainability Forecast
Are your best ideas showing?
We flag the genuinely novel ideas in your codebase — algorithms, architectures, or methods worth protecting — and check two separate ways they can leak: through a public repository, and through public-facing copy (README, docs, marketing) that spells out how something works rather than just what it does. This is an independent signal and never affects your production verdict.
module ip_exposure · priority tier Specialized · Claude-reasoned · findings land under IP & Novelty Exposure
And if you handle personal data: GDPR readiness
A focused pass on GDPR readiness: whether you have a lawful basis for the personal data you collect, real consent flows, a clear picture of where personal data lives, a working path to honor deletion and data-subject-rights requests, and a plan for cross-border transfers and breach notification. Findings here are grouped with the broader Privacy & Compliance check. nittim runs this automatically when a repository shows signs of handling personal data.
How often we’re wrong
false criticals — clean, mature libraries, both tiers
intentionally-vulnerable apps verdicted unsafe
secret false-criticals — real, deployed apps
mature deployed apps passed green, Deep tier
Small corpora, honest caveats — every number ships with its method and confidence interval.
0 of 25 false criticals → 95% CI 0–13.8% (Clopper-Pearson). Verdict-level recall 100% (11 of 11); 91.7% under strict pre-registration. The libraries corpus is IN-SAMPLE — four precision rules were written against its failures — so read it beside the 12-library held-out number (1 of 12 secret FCR, pre-registered before it ran), never alone. The judge shares the Opus family; only the separate MCP cross-check is cross-vendor. Every not-green deployed-app verdict was adjudicated over-escalated on dependency drift; root causes fixed, re-run outstanding.
A human audit runs $500–$3,000 and takes days to weeks. nittim deep-verifies a repo in minutes — evidence included.
Is my code safe with you?
You shouldn't have to trust nittim before you get anything from it. Start at the top.
0 · Your assistant, our rubric
Run /selfcheck’s 13 categories with your own assistant — you give us nothing. It wrote the code, so it may repeat its own blind spots; an independent judge catches those, and that’s what we sell.
1 · Our report library
Read real scans of well-known public code in /library before nittim sees a line of yours. Nothing to take on faith.
2 · Your repo, free scan
No AI in the free scan. Two deterministic scanners read the files and build the report. A test fails the build if source text reaches the database — a property of the code, not a promise.
3 · AI audit on your own key
BYOK Pro — API, MCP, or GitHub Action: the call runs on your key, under your vendor terms. nittim's model account never sees your code — check your own Anthropic bill.
4 · AI audit, on our key
Held in memory for the audit, never trained on, discarded after. The one rung that asks you to take nittim's word for it — last on purpose.
5 · Your own infrastructure
Built, both halves ran end to end, and it runs in your account — no source, finding or repo name reaches nittim. Not self-serve: we set it up with you on request. Images aren’t signed yet — verify by published hash.
- Stored
- Audit report — findings, scores, verdict, repository name.
- Not stored
- Source code — fetched into server memory during analysis, discarded when the pipeline finishes. Exception: finding evidence includes the lines it points to.
- Not stored
- GitHub tokens — used once, discarded. Installation tokens expire within the hour; nittim keeps only the installation id.
- Read-only
- GitHub App: exactly two permissions — contents and metadata, both read-only — no events. nittim can never modify your code.
What it costs
nittim never bills flat rate for tokens it doesn't control — prepaid credits, your own Anthropic key, or your own cloud, your choice, never a surprise.
- Free
- Deterministic secret + dependency-CVE scan — any public or connected repo, forever. No card, no sign-in. One covered Audit run per account (GitHub or Google).
- Credits
- 1 credit = $1. An AI audit costs 5.14 credits; a single module via API/MCP costs 5.03. Packs: $29 → 5 · $74 → 17 · $189 → 52 · $479 → 157 audits. Credits never expire · unused credits refundable within 30 days.
- BYOK Pro
- $49/mo pays for the pipeline, not tokens — bring your own Anthropic key. A 100-credit monthly allowance covers ~50 single-pass audits at 2 credits each.
- Enterprise
- from $18k/yr — seat bands, unlimited repos, SSO, SARIF, policy-as-code, append-only audit log. Runs on Claude via Bedrock in your AWS, or your own Anthropic org. Orgs share one audit portfolio and one bill. Talk to us →