Francisco MinguezData & AI

Services

Technical consulting for AI and data systems that need evidence.

I take on focused engagements when a team needs to evaluate reliability, compare models or configurations, protect a release, or improve the system underneath, with evidence kept explicit before the next commercial move.

Situations

Before shipping, when choosing, or after a change.

Each situation maps to a different kind of evidence work.

  1. 01

    Before shipping: is it reliable enough?

    Independent, time-boxed evaluation of an existing RAG system for reliability and go-live risk, plus prioritised remediation.

    RAG Reliability Audit

  2. 02

    When choosing: which model or configuration makes sense?

    Frozen synthetic comparison of candidates. Not RAG-specific. A valid outcome is NO ELIGIBLE RECOMMENDATION.

    LLM Model Selection Benchmark

  3. 03

    After a change: did we make it worse?

    Fixture-first release gate: baseline PASS, then mutation, then critical FAIL, then release BLOCKED. Not a model-selection product.

    GenAI Regression Gate

Readiness evaluation

Working on your system, step by step.

A readiness evaluation looks at your system under agreed access. Public Study and Explorer metrics are never treated as your customer results.

  1. 01

    Understand the risk

    Clarify the decision (pilot, go-live, handover) and which reliability questions matter.

  2. 02

    Connect and check access

    Confirm runnable access (or an agreed equivalent), representative documents or examples, and a named technical counterpart.

  3. 03

    Freeze the evaluation contract

    Agree scope, dimensions, assumptions, deliverables and exclusions. This is not certification, compliance sign-off, or a production guarantee.

  4. 04

    Evaluate

    Time-boxed independent evaluation on the client system within the frozen contract.

  5. 05

    Decide with evidence

    Executive decision summary, dimension scorecard, critical failure modes, prioritised remediation and a readout session.

  6. 06

    Protect later changes when relevant

    If prompts, retrieval or models will keep changing, we can connect to regression-protection thinking. Optional, and not required for every Audit.

The point is to reduce the chance of shipping or expanding on anecdote (rework, stalled handovers, eroded trust). No invented ROI percentage or guaranteed savings.

Ways of working

  1. 01

    Technical evaluation

    I review a data, ML or RAG system and provide findings and practical next decisions without requiring an implementation phase.

  2. 02

    Defined project

    I take ownership of a bounded technical outcome with agreed deliverables and handover.

  3. 03

    Embedded support

    I join an existing team part-time or for a fixed period, directly or through a consultancy. Useful when you need capacity; secondary to specialised evaluation and delivery offers.

Technical training and knowledge transfer can support an engagement when they help your team operate or extend the work.

Engagement detail

What I do, what you receive, and when it does not fit.

The RAG Reliability Audit is the packaged evaluation for production-readiness questions on an existing RAG system. Pipeline modernisation and applied ML remain available when those problems are the real need.

Production readiness evaluation

RAG Reliability Audit

I run a time-boxed independent evaluation of an existing RAG system so you get bounded decision support on reliability and go-live risk, plus prioritised remediation. It is not certification or a production guarantee.

This may help when

  • A RAG system looks convincing in demos, but go-live, pilot expansion or client handover still rests on anecdote.
  • Retrieval quality, grounded answering and abstention when evidence is thin are not evaluated with explicit criteria.
  • The team sees failures but lacks a representative evaluation set and a practical remediation order.

What I can deliver

  • Executive decision summary: reliability and go-live-risk assessment for a go / no-go conversation.
  • Dimension scorecard covering retrieval, answer quality, abstention, failure modes and regression readiness within agreed scope.
  • Critical failure modes observed or strongly indicated, with assumptions and limitations stated.
  • Prioritised remediation list and one clarification / readout session.

What I need to start

  • You have runnable access to the RAG system under review, or an agreed equivalent environment.
  • You can provide representative documents or examples and a clear decision question (pilot, go-live or handover).
  • A named technical counterpart can clarify access, scope and failure concerns.

Specific limits

  • You need a guarantee that the system is production-ready, or guaranteed answer quality without representative evidence.
  • You need compliance certification, legal sign-off or a security audit.

Related synthetic study

I built the public RAG Reliability Study v1 to test a pre-registered answer/defer gate on a frozen synthetic holdout. It is methodological evidence on a synthetic benchmark, not proof of production readiness and never your customer results. The separate RAG Reliability Explorer is a synthetic demo runtime, not a customer runtime. An engagement evaluates your system under agreed access.

View related case study

Related technical work

Useful when the problem fits. Described from professional experience, not as separately packaged commercial products.

Data pipeline modernisation and quality

I migrate or improve recurring analytical workflows so their transformations, validation controls and operating responsibilities are easier to understand and maintain.

This may help when

  • A recurring SQL, SAS or ETL process is difficult to maintain or change safely.
  • Data-quality problems are detected late or handled manually.
  • Legacy and target results are difficult to compare and reconcile.

What I can deliver

  • Review of the current workflow, dependencies and business rules.
  • A staged migration and reconciliation plan.
  • Reusable transformations with automated data-quality and validation controls.
  • Operational documentation and practical handover.

What I need to start

  • A concrete workflow and expected outputs already exist.
  • Legacy and target results can be compared.
  • Your team can provide the required data access and technical ownership.

Specific limits

  • You need a full platform replacement before the current workflow is understood.
  • No owner can explain the existing rules or expected outputs.

Relevant professional experience

I worked on the migration of a recurring analytical process from legacy SQL Server logic to a modular Databricks and PySpark workflow with reusable components, data-quality controls and validation.

View related case study

Applied ML and decision workflows

I turn an existing operational decision into a testable analytical workflow: establish a baseline, prepare the available data, evaluate errors and deliver outputs that people can use in the process.

This may help when

  • A prediction, prioritisation or document-analysis problem lacks a clear baseline.
  • Manual analytical or document review is difficult to scale.
  • Model outputs are disconnected from the decision people need to make.

What I can deliver

  • Problem framing and baseline definition.
  • Data preparation and diagnostic analysis.
  • A predictive, NLP or analytical workflow with error analysis.
  • Usable outputs, documentation and knowledge transfer.

What I need to start

  • A concrete decision or operational process can be described.
  • Representative historical or labelled data is available.
  • We can agree a baseline and a useful evaluation criterion.

Specific limits

  • You need a model before the decision problem is defined.
  • You require guaranteed predictive or business performance.

Relevant professional experience

My professional experience includes applied analytics, predictive modelling, NLP, reporting automation and communication with technical and business stakeholders. Organisation-specific examples and artefacts remain confidential.

Model choice and regression protection

Synthetic evidence for two adjacent decisions.

These support model or configuration choice and release protection. They are controlled demos, not client results.

Are we using the right model or configuration?

LLM Model Selection Benchmark

Compares model or configuration candidates under a frozen synthetic evaluation so teams can see whether any option clears an agreed quality bar.

  • Synthetic benchmark evidence.
  • Observed outcome: NO ELIGIBLE RECOMMENDATION (valid technical stop).
  • Holdout: CONSUMED. Future unseen validation needs a new holdout.
  • Not RAG-specific positioning.

Controlled synthetic demo. Not a client result or a production model recommendation.

View evidence summary

How do we stop future changes from breaking it?

GenAI Regression Gate

Shows a release gate that blocks a change introducing a critical behaviour regression.

  • SYNTHETIC REGRESSION GATE DEMO (fixture-first).
  • Causal chain: baseline PASS, then prompt mutation, then critical regression, then FAIL, then release BLOCKED.
  • Not a model-selection product.

Fixture-first synthetic demo. Not live GenAI causal proof or a production guarantee.

View evidence summary

Supporting research

Study and Explorer explain the method.

Useful to inspect how the work thinks. They are not packaged offers.

Research / methodology evidence

RAG Reliability Study

Public synthetic study of answer/defer behaviour on a frozen holdout. Methodological evidence only; not proof of production readiness.

Interactive synthetic demo

RAG Reliability Explorer

Separate demo runtime with its own synthetic examples. Not a customer runtime and not a replay of study metrics.

What helps before a first conversation

  • What exists today: a workflow, system, dataset or prototype.
  • What you need to improve, understand or decide.
  • The representative data, documents or examples available.
  • Any important timing, security or confidentiality constraints.

What we will confirm before starting

Before agreeing the work, we will confirm the available evidence, the delivery boundary and what a useful outcome can realistically mean.

  • The scope, deliverables and acceptance criteria.
  • What the available data and access allow us to evaluate.
  • Which production, security or compliance work falls outside the engagement.
  • Which results can be measured without treating them as guarantees.

Managed business solutions

Do you need a managed business solution rather than individual consultancy?

I also founded Miga Digital, an AI and automation studio for small businesses in Spain. If you need a solution to be implemented, managed and improved as a service, I can help determine whether Miga Digital is the better route. Visit Miga Digital

Tell me what you're building.

Describe the system or workflow, what you are unsure about, and what a bad outcome would look like. We can decide together whether evaluation, model comparison, regression protection or broader delivery work is the right next step.

Tell me what you're building