OwlShield
A firewall that sits between applications and an LLM, screening prompts for attacks and responses for harmful output.
- Year
- 2024
- Team
- Four, built alongside national service
- My part
- Full-stack developer
- Stack
- Svelte and Supabase, with Python for evaluation
- Result
- 1st place, NSCC HPC Innovation Competition 2023, university category
The problem
Teams shipping LLM features have little between the user and the model. Prompt injection and jailbreaks go straight in; unsafe or leaked content comes straight out. Most mitigations are baked into each application, so every team rebuilds them.
How it works
OwlShield runs as a proxy in front of the model. Each request passes through input checks before it reaches the LLM, and each response through output checks before it reaches the user. Policies are configured once and apply to every app behind the proxy, which gives teams one place to keep their apps safe and in line with AI regulations.
Testing it
To measure what OwlShield caught, we wanted to run Garak, NVIDIA's open-source LLM vulnerability scanner, against it. I wrote a LiteLLM integration for Garak so its probes could reach OwlShield, and contributed it upstream; it was merged in April 2024.
The competition
We built OwlShield for the HPC Innovation Competition 2023, run by Singapore's National Supercomputing Centre (NSCC). The work ran from January to July 2024, and the team won first place in the university category.