Back

ResearchJul 2025 to Oct 2025

Malicious prompt classifier

Stop harmful prompts before they reach a language model, and say why.

  • Probabilistic modelling
  • Markov chains
  • AI safety
  • Python

Impact

precision
98.84%
accuracy
90.79%
F1 score
89.96%

Published in Procedia Computer Science, 2026.

What I did

  • Developed a probabilistic detector that scores a prompt by the transitions between its text sequences, modelled as a Markov chain.
  • Added an explanation module that highlights the high-risk patterns behind each flag.
  • Reached 90.79% accuracy, 98.84% precision, 82.54% recall and an 89.96% F1 score on a malicious-prompt benchmark.
  • Published the paper in Procedia Computer Science (2026).

The aim

Chatbots can be tricked by prompts written to make them misbehave. This research catches those prompts before they get through, and says why each one was flagged.