ResearchJul 2025 to Oct 2025
Malicious prompt classifier
Stop harmful prompts before they reach a language model, and say why.
- Probabilistic modelling
- Markov chains
- AI safety
- Python
Impact
- precision
- 98.84%
- accuracy
- 90.79%
- F1 score
- 89.96%
Published in Procedia Computer Science, 2026.
What I did
- Developed a probabilistic detector that scores a prompt by the transitions between its text sequences, modelled as a Markov chain.
- Added an explanation module that highlights the high-risk patterns behind each flag.
- Reached 90.79% accuracy, 98.84% precision, 82.54% recall and an 89.96% F1 score on a malicious-prompt benchmark.
- Published the paper in Procedia Computer Science (2026).
The aim
Chatbots can be tricked by prompts written to make them misbehave. This research catches those prompts before they get through, and says why each one was flagged.