Francisco MinguezData & AI

Selected work

Selected professional work.

Applied AI, machine learning and data engineering in delivery contexts I can stand behind, with evidence labels kept explicit. The Evidence Atlas remains the broader corpus of research, benchmarks, demos and methodology.

Decision evidence

I prepare artefacts for readiness, model or configuration choice and regression protection, plus a subordinate interactive demo when you want to poke the behaviour.

Reproducible RAG evaluation study

RAG Reliability study

Research evidence: I built this study to test when a RAG system should answer or defer using a versioned synthetic benchmark and a fixed evaluation protocol. It is not client results and not a packaged commercial offer.

  • Python
  • uv
  • pytest
  • BM25
  • sentence-transformers
  • all-MiniLM-L6-v2
  • reciprocal-rank fusion
  • SHA-256 artifact verification

Read case study

Synthetic benchmark

LLM Model Selection Benchmark

Compares model or configuration candidates under a frozen synthetic evaluation so a team can decide whether any option clears an agreed quality bar. Not a RAG-specific product.

Observed synthetic outcome: NO ELIGIBLE RECOMMENDATION. Holdout: CONSUMED.

Synthetic benchmark only. Controlled demo, not a client result or production recommendation.

Synthetic gate demo

GenAI Regression Gate

Shows how a release gate can block a change that introduces a critical behaviour regression: baseline PASS, one prompt mutation, critical failure, candidate FAIL, release BLOCKED.

Evidence class: SYNTHETIC REGRESSION GATE DEMO (fixture-first).

Fixture-first synthetic demo. Not live GenAI causal proof or a model-selection product.

Interactive synthetic demo

RAG Reliability Explorer

A subordinate interactive demo with its own synthetic examples. It does not replay the study benchmark, does not produce client metrics, and is not a packaged offer.

Open the live demo (opens in a new tab)

Data Engineering

I modernise data workflows, improve quality controls and make analytical processes easier to validate, operate and hand over.

Professional data engineering delivery

Analytical pipeline modernisation

I worked on the migration of a recurring analytical process from legacy SQL Server logic to a modular Databricks and PySpark workflow with data-quality and validation controls.

  • PySpark
  • Databricks
  • SQL Server
  • SQL
  • Data quality
  • Statistical validation

Read case study

Public synthetic implementation

Data Quality Pipeline Lab

I built a local PySpark reference pipeline with explicit contracts, quarantine, reconciliation, idempotent runs, tests and CI on fully synthetic B2B wholesale data. It does not reproduce an employer's code, data or systems.

  • Python
  • PySpark
  • uv
  • pytest
  • Ruff
  • mypy
  • GitHub Actions
  • Parquet

Read case study

How to read this work

  • Decision artefacts answer ship, model or regression questions; supporting data work is separate.
  • Labels distinguish professional experience, research studies, synthetic demos and interactive support demos.
  • Study and Explorer support RAG reliability work; they are not peer packaged offers.
  • A public demo or synthetic benchmark illustrates a method. It is not evidence of production performance.