I test how untrusted text becomes an instruction, a tool call, or a claim people rely on. Then I build the controls that should have stopped it.
11Case studies
26Public evidence records
4Merged upstream PRs
The evidence
Claims you can inspect.
Original captures, selected source records, and clear scope. Expand a case, then its evidence.
Severity and CVSS are researcher assessments. Vendor responses and evidence limits are recorded per case. Studies and observations are excluded from vulnerability counts.
Featured case · Jack & Jill AI
An offer I never applied for. Then an email to prove it.
A hiring agent gave me a Founding Engineer offer I never applied for — then put it in writing over email.
7C / 4H / 3M14 original reported findings
Original vulnerability confirmed by the founder · May 2026
The £100,000 offer
The agent presents an offer for a role I had not applied to. This is a chat claim, not a real employment contract.
May 2026 · 01-jack-announces-100k-offer.pngThe claim leaves the chat
An email received from jack@jackandjill.ai repeats the fabricated terms. The email delivery is real. The job and sponsorship claims are not validated.
Boundary testing and runtime enforcement for deployed AI, 2026–present. In development.
IntentScan generates adversarial probes against any AI API — role transformation, gradual drift, language variation — and scores capability, role, and domain violations through an LLM judge into a 0–100 risk report
IntentEnforce is a runtime proxy that classifies user intent per request and applies allow / block / clarify policy before traffic reaches the model
Python and FastAPI, React and TypeScript, on Firebase
Adversarial benchmarking
AgentInjectionBench
Open benchmark (Apache-2.0, 2026) for prompt injection against agentic tool-use and MCP-style integrations — the surface that only exists once a model can call tools.
SSRN, Jun 2026, DOI 10.2139/ssrn.6874522, sole author. A severity-weighted refusal metric, gaming-resistant by proof. The pilot showed flat averaging hides a prompt-injection weakness in Llama 3.3 70B — 0.800 flat against 0.730 WSR.
A red-teaming playground I built and open-sourced (ppradyoth/prompt-injection-ctf): 16 challenges covering all 10 OWASP LLM Top 10 (2025) risks, with a defender mode that reveals the guardrail code behind each attack. Unprompted, an AI engineer at Apiiro (Shmulik Cohen) froze 11 of its system prompts as the fixtures for a nine-model experiment in August 2026 — 18 attacks, three runs each, 486 attempts, published with code and data and crediting the CTF by name.
GPT-3.5 fell 54/54; GPT-5.6 Sol dropped to 6/54 — real progress
But two of the CTF's attacks still landed on every attempt against the newest model, with poisoned documents driving SQL and command injection
Newer models learned to rank what they trust; they still have no wall between instructions and data
Weighted Safety Refusal (WSR): A Reference-free, Severity-weighted, Dual-axis Metric for Evaluating LLM Refusal Behavior
SSRN preprint · June 2026 · Sole author
Flat refusal averages hide the failures that matter. WSR weights refusals by severity across two axes and is gaming-resistant by proof — in the pilot, Llama 3.3 70B scores 0.800 flat but 0.730 under WSR, and the gap is prompt injection.
Analyzing the Difficulties in Major Applications of Augmented Reality
BIBLUS — National Level Paper Presentation, NIE IEEE Student Branch · 2021 · II prize
The Prompt Injection CTF was adopted by an Apiiro engineer for a published 9-model study. Credential Guard is an upstream proposal. The original Jack & Jill vulnerability was confirmed by its founder.
All work here was performed independently, on personal time, equipment, and accounts, outside and unrelated to employment. It uses no employer systems, data, or resources and does not represent any employer's views. Disclosure history and evidence limits are recorded per case. Public exhibits omit attack payloads and sensitive source material. Full reproduction is available on request.