I prepare artefacts for readiness, model or configuration choice and regression protection, plus a subordinate interactive demo when you want to poke the behaviour.
Reproducible RAG evaluation study
RAG Reliability study
Research evidence: I built this study to test when a RAG system should answer or defer using a versioned synthetic benchmark and a fixed evaluation protocol. It is not client results and not a packaged commercial offer.
- Python
- uv
- pytest
- BM25
- sentence-transformers
- all-MiniLM-L6-v2
- reciprocal-rank fusion
- SHA-256 artifact verification
Read case study
Synthetic benchmark
LLM Model Selection Benchmark
Compares model or configuration candidates under a frozen synthetic evaluation so a team can decide whether any option clears an agreed quality bar. Not a RAG-specific product.
Observed synthetic outcome: NO ELIGIBLE RECOMMENDATION. Holdout: CONSUMED.
Synthetic benchmark only. Controlled demo, not a client result or production recommendation.
Synthetic gate demo
GenAI Regression Gate
Shows how a release gate can block a change that introduces a critical behaviour regression: baseline PASS, one prompt mutation, critical failure, candidate FAIL, release BLOCKED.
Evidence class: SYNTHETIC REGRESSION GATE DEMO (fixture-first).
Fixture-first synthetic demo. Not live GenAI causal proof or a model-selection product.
Interactive synthetic demo
RAG Reliability Explorer
A subordinate interactive demo with its own synthetic examples. It does not replay the study benchmark, does not produce client metrics, and is not a packaged offer.
Open the live demo (opens in a new tab)