Independent AI security research

Where AI systems
lose the boundary.

I test how untrusted text becomes an instruction, a tool call, or a claim people rely on. Then I build the controls that should have stopped it.

11Case studies
26Public evidence records
4Merged upstream PRs

The evidence

Claims you can inspect.

Original captures, selected source records, and clear scope. Expand a case, then its evidence.

Severity and CVSS are researcher assessments. Vendor responses and evidence limits are recorded per case. Studies and observations are excluded from vulnerability counts.

Featured case · Jack & Jill AI

An offer I never applied for.
Then an email to prove it.

A hiring agent gave me a Founding Engineer offer I never applied for — then put it in writing over email.

7C / 4H / 3M14 original reported findings

Original vulnerability confirmed by the founder · May 2026

The £100,000 offer

The agent presents an offer for a role I had not applied to. This is a chat claim, not a real employment contract.

May 2026 · 01-jack-announces-100k-offer.png
The claim leaves the chat

An email received from jack@jackandjill.ai repeats the fabricated terms. The email delivery is real. The job and sponsorship claims are not validated.

May 28, 2026 · Offer email capture
Full transcripts and reproduction on request.ppradyoth64@gmail.com
Reconnaissance & research notes2 packages

Google AI

Attack planning, program-scope notes, and references to previously disclosed work. No independently confirmed finding is presented from this package.

LegalOS

Intake and reconnaissance notes. No demonstrated exploit or security-impact claim is presented from this package.

AI Security Work

Contributing to the tools
that secure AI.

I contribute fixes to NVIDIA garak, Promptfoo, and Presidio, and build tools for adversarial testing, credential protection, and risk evaluation.

Open-source contributions

4 merged8 open

Status checked · View on GitHub ↗

More upstream work

Security tooling & evaluation

Runtime enforcement

Akrivon

Boundary testing and runtime enforcement for deployed AI, 2026–present. In development.

  • IntentScan generates adversarial probes against any AI API — role transformation, gradual drift, language variation — and scores capability, role, and domain violations through an LLM judge into a 0–100 risk report
  • IntentEnforce is a runtime proxy that classifies user intent per request and applies allow / block / clarify policy before traffic reaches the model
  • Python and FastAPI, React and TypeScript, on Firebase

Adversarial benchmarking

AgentInjectionBench

Open benchmark (Apache-2.0, 2026) for prompt injection against agentic tool-use and MCP-style integrations — the surface that only exists once a model can call tools.

Security measurement

Weighted Safety Refusal

SSRN, Jun 2026, DOI 10.2139/ssrn.6874522, sole author. A severity-weighted refusal metric, gaming-resistant by proof. The pilot showed flat averaging hides a prompt-injection weakness in Llama 3.3 70B — 0.800 flat against 0.730 WSR.

Security education

Prompt Injection CTF

A red-teaming playground I built and open-sourced (ppradyoth/prompt-injection-ctf): 16 challenges covering all 10 OWASP LLM Top 10 (2025) risks, with a defender mode that reveals the guardrail code behind each attack. Unprompted, an AI engineer at Apiiro (Shmulik Cohen) froze 11 of its system prompts as the fixtures for a nine-model experiment in August 2026 — 18 attacks, three runs each, 486 attempts, published with code and data and crediting the CTF by name.

  • GPT-3.5 fell 54/54; GPT-5.6 Sol dropped to 6/54 — real progress
  • But two of the CTF's attacks still landed on every attempt against the newest model, with poisoned documents driving SQL and command injection
  • Newer models learned to rank what they trust; they still have no wall between instructions and data

Publications & independent use

Research beyond the test.

Weighted Safety Refusal (WSR): A Reference-free, Severity-weighted, Dual-axis Metric for Evaluating LLM Refusal Behavior

SSRN preprint · June 2026 · Sole author

Flat refusal averages hide the failures that matter. WSR weights refusals by severity across two axes and is gaming-resistant by proof — in the pilot, Llama 3.3 70B scores 0.800 flat but 0.730 under WSR, and the gap is prompt injection.

Analyzing the Difficulties in Major Applications of Augmented Reality

BIBLUS — National Level Paper Presentation, NIE IEEE Student Branch · 2021 · II prize

The Prompt Injection CTF was adopted by an Apiiro engineer for a published 9-model study. Credential Guard is an upstream proposal. The original Jack & Jill vulnerability was confirmed by its founder.

All work here was performed independently, on personal time, equipment, and accounts, outside and unrelated to employment. It uses no employer systems, data, or resources and does not represent any employer's views. Disclosure history and evidence limits are recorded per case. Public exhibits omit attack payloads and sensitive source material. Full reproduction is available on request.