Francisco MinguezData & AI

Verified professional work

Published case study

Analytical pipeline modernisation

Migration of a recurring analytical process from a legacy SQL Server implementation to a modular Databricks and PySpark workflow with automated quality controls.

Organisation-specific details are omitted. The page documents verified responsibilities and reported outcomes without publishing confidential data or internal artefacts.

Context

A recurring analytical process in a financial-data environment. Organisation, dataset and internal implementation details are intentionally omitted.

Problem

The existing calculation depended on a legacy SQL Server workflow and required weeks to complete, making validation, iteration and operational delivery slow.

Scope

  • Map the legacy calculation logic and expected outputs.
  • Reimplement the workflow with Databricks and PySpark.
  • Create reusable, parameterised transformation components.
  • Add automated data-quality controls and statistical validation.
  • Support technical integration, documentation and handover.

Constraints

  • Existing business rules had to remain traceable while the implementation changed.
  • Differences between legacy and migrated outputs required explicit reconciliation.
  • Organisation-specific data, code and architecture cannot be published.
  • The public case study must describe the delivery without exposing confidential information.

Approach

  1. 01

    Frame the migration

    Document the calculation rules, dependencies, expected outputs and acceptance criteria before changing the implementation.

  2. 02

    Rebuild the workflow

    Translate the process into reusable and parameterised PySpark components running on Databricks.

  3. 03

    Validate the outputs

    Apply automated data-quality checks, statistical validation and output reconciliation against the legacy workflow.

  4. 04

    Prepare operational handover

    Document execution requirements, known limitations and integration considerations for the teams operating the process.

Deliverables

  • Migrated Databricks and PySpark workflow.
  • Reusable parameterised transformation components.
  • Automated data-quality controls.
  • Statistical validation and reconciliation workflow.
  • Technical documentation and handover material.

Technologies

  • PySpark
  • Databricks
  • SQL Server
  • SQL
  • Data quality
  • Statistical validation

Evidence

  • Verified professional result

    Calculation cycle

    Weeks to approximately two hours

    Reported from the delivered professional workflow. Internal timing records and organisation-specific validation artefacts are not public.

  • Verified implementation evidence

    Engineering structure

    Reusable parameterised PySpark components

    Components incorporated automated data-quality controls and statistical validation.

Limitations

  • The underlying code, data and organisation-specific architecture are confidential.
  • The public page cannot independently reproduce the original working environment.
  • The reported runtime is specific to the original workload, infrastructure and implementation.
  • No claim is made that the same runtime improvement would apply to another organisation without assessment.

Production considerations

  • Orchestration, retries and failure recovery.
  • Data contracts and schema-change handling.
  • Access control and sensitive-data governance.
  • Monitoring, lineage and operational alerting.
  • Compute-cost and performance tuning.
  • Reconciliation, rollback and controlled migration.

Service relevance

  • SQL or SAS analytical-process migration to PySpark.
  • Databricks pipeline and workflow modernisation.
  • ETL and data-quality design.
  • Analytical-process automation.
  • Technical audit and migration roadmap.
  • Delivery documentation and handover.

Next steps

Back to selected work