Case Study 02PharmaceuticalsAWS / Multi-Cloud

    Databricks Pipeline Optimization with Infrastructure as Code

    AstraZeneca, a global pharmaceutical and bioscience company, was facing significant challenges with inconsistent deployment processes, manual Hive metastore management, and fragmented data ingestion from multiple sources. The lack of infrastructure automation led to environment drift, slow CI/CD cycles, and difficulty replicating jobs across environments. Vrahad Analytics was brought in to implement a comprehensive Infrastructure as Code strategy that would transform AstraZeneca's data engineering operations into a modern, repeatable, and scalable platform.

    Client

    AstraZeneca Pharma

    Cloud Platform

    AWS / Multi-Cloud

    Duration

    5 Months

    Databricks Pipeline Optimization with Infrastructure as Code

    The Challenge

    AstraZeneca's data engineering team had built their Databricks environment manually over time, resulting in configuration drift between development, staging, and production environments. This created a host of reliability and productivity issues that compounded as the team grew.

    1

    Manual Databricks workspace provisioning and job configuration led to significant environment drift between dev, staging, and production, causing unpredictable failures during deployments.

    2

    Hive metastore tables lacked centralized governance — schema changes in one environment didn't propagate to others, creating data inconsistencies and breaking downstream consumers.

    3

    Data ingestion from Salesforce and other SaaS platforms was handled through custom scripts that required constant maintenance, with no monitoring or automated recovery mechanisms.

    4

    CI/CD pipelines were rudimentary, relying on manual deployment steps that slowed the release cycle from days to weeks and introduced human error into the deployment process.

    5

    No version control for infrastructure configurations meant that recovering from misconfigurations required manual investigation and reconstruction, increasing mean time to recovery.

    6

    Cross-team collaboration was hampered by the inability to quickly spin up isolated development environments that mirrored production configurations.

    Our Solution

    We implemented a comprehensive Infrastructure as Code strategy using Terraform as the backbone, enabling consistent and repeatable Databricks deployments across all environments. The project encompassed infrastructure automation, data ingestion modernization, and CI/CD pipeline establishment within a modern Lakehouse architecture.

    Designed and implemented Terraform modules for the entire Databricks infrastructure, including workspace configuration, cluster policies, job definitions, secret scopes, and Unity Catalog objects — all version-controlled in Git.

    Migrated all Hive metastore tables to Unity Catalog with automated schema validation, enabling centralized governance, cross-workspace data sharing, and complete audit trail capabilities.

    Deployed Fivetran connectors for automated, reliable data ingestion from Salesforce, replacing fragile custom scripts with a managed service that provides schema drift detection, automatic retry logic, and data freshness monitoring.

    Established robust CI/CD pipelines using GitHub Actions integrated with Terraform, enabling infrastructure changes to go through automated plan → review → apply workflows with approval gates.

    Built a Lakehouse architecture that organizes data into bronze (raw), silver (cleaned), and gold (business-ready) layers, with automated data quality checks at each transformation stage.

    Created self-service Terraform templates that enable AstraZeneca's engineering teams to provision new data pipeline environments in minutes instead of days, with guardrails enforced through policy-as-code.

    Implementation Phases

    1

    Current State Assessment

    3 Weeks

    Audited all existing Databricks configurations, Hive metastore schemas, custom ingestion scripts, and deployment processes. Documented the gap between current state and target architecture, and prioritized remediation items.

    2

    Terraform Foundation

    4 Weeks

    Developed core Terraform modules for Databricks workspace provisioning, cluster policy management, job definitions, and Unity Catalog objects. Established the remote state backend, module registry, and CI/CD integration.

    3

    Hive to Unity Catalog Migration

    4 Weeks

    Migrated all Hive metastore tables to Unity Catalog in a phased approach, validating data integrity at each step. Implemented automated schema synchronization and access control policies across all environments.

    4

    Fivetran Integration & Data Ingestion

    3 Weeks

    Deployed Fivetran connectors for Salesforce and other data sources, configured destination schemas in Unity Catalog, set up freshness monitoring, and decommissioned legacy custom ingestion scripts.

    5

    CI/CD Pipeline Implementation

    3 Weeks

    Built end-to-end CI/CD pipelines for both infrastructure (Terraform) and data pipeline (Databricks jobs) deployments. Implemented automated testing, approval gates, and rollback mechanisms.

    6

    Validation & Knowledge Transfer

    3 Weeks

    Performed comprehensive end-to-end testing across all environments, documented all Terraform modules and CI/CD processes, and conducted hands-on training sessions for AstraZeneca's engineering team.

    Technologies Used

    TerraformDatabricks Unity CatalogFivetranHive MetastoreSalesforceCI/CD PipelinesLakehouse ArchitectureGitHub ActionsDelta LakePySparkPythonAWS

    Key Results

    Measurable outcomes and business impact delivered through this engagement.

    Achieved consistent, repeatable deployments across all Databricks environments (dev, staging, production) with zero configuration drift, reducing deployment-related incidents by 90%.

    Eliminated environment drift entirely through Terraform-managed infrastructure, ensuring that every environment is an exact replica of the declared configuration in version control.

    Automated data ingestion from Salesforce and other sources via Fivetran, reducing ingestion-related maintenance effort by 80% and improving data freshness from daily to near real-time.

    Enhanced data governance capabilities with Unity Catalog's built-in lineage tracking, audit logging, and granular access controls, meeting pharmaceutical regulatory requirements.

    Accelerated CI/CD cycles for data pipeline deployment from weeks to hours, with automated testing and approval workflows that maintain quality while increasing velocity.

    Enabled self-service environment provisioning, allowing teams to spin up new development environments in under 15 minutes with production-parity configurations.

    Before vs. After Comparison

    Deployment Time

    Before

    2-3 weeks

    After

    2-3 hours

    95% faster

    Environment Drift

    Before

    Frequent issues

    After

    Zero drift

    100% eliminated

    Data Freshness

    Before

    Daily batches

    After

    Near real-time

    24x improvement

    Ingestion Maintenance

    Before

    20+ hrs/week

    After

    4 hrs/week

    80% reduction

    Env Provisioning

    Before

    3-5 days

    After

    15 minutes

    99% faster

    Deployment Incidents

    Before

    8-10/month

    After

    < 1/month

    90% reduction

    Key Infrastructure Metrics (Post-Implementation)

    Env Provision (min)Terraform Modules07142128

    Effort Distribution by Workstream

    Platform Maturity — Before vs After

    RepeatabilityAutomationGovernanceVelocityReliabilitySelf-Service0255075100
    Before After

    Infrastructure Automation Coverage (%)

    Month 1Month 2Month 3Month 4Month 50255075100