Technical area / 01
AI and LLM evaluation
Relevance, consistency, behaviour, and safety
I use structured evaluation practices to examine generated output, agent behaviour, retrieval quality, and natural-language workflows.
- AGolden datasets and prompt variation
- BSemantic similarity and response evaluation
- CReadable quality signals