Databricks Pipeline Optimization with Infrastructure as Code
AstraZeneca, a global pharmaceutical and bioscience company, was facing significant challenges with inconsistent deployment processes, manual Hive metastore management, and fragmented data ingestion from multiple sources. The lack of infrastructure automation led to environment drift, slow CI/CD cycles, and difficulty replicating jobs across environments. Vrahad Analytics was brought in to implement a comprehensive Infrastructure as Code strategy that would transform AstraZeneca's data engineering operations into a modern, repeatable, and scalable platform.
Client
AstraZeneca Pharma
Cloud Platform
AWS / Multi-Cloud
Duration
5 Months
The Challenge
AstraZeneca's data engineering team had built their Databricks environment manually over time, resulting in configuration drift between development, staging, and production environments. This created a host of reliability and productivity issues that compounded as the team grew.
Manual Databricks workspace provisioning and job configuration led to significant environment drift between dev, staging, and production, causing unpredictable failures during deployments.
Hive metastore tables lacked centralized governance — schema changes in one environment didn't propagate to others, creating data inconsistencies and breaking downstream consumers.
Data ingestion from Salesforce and other SaaS platforms was handled through custom scripts that required constant maintenance, with no monitoring or automated recovery mechanisms.
CI/CD pipelines were rudimentary, relying on manual deployment steps that slowed the release cycle from days to weeks and introduced human error into the deployment process.
No version control for infrastructure configurations meant that recovering from misconfigurations required manual investigation and reconstruction, increasing mean time to recovery.
Cross-team collaboration was hampered by the inability to quickly spin up isolated development environments that mirrored production configurations.
Our Solution
We implemented a comprehensive Infrastructure as Code strategy using Terraform as the backbone, enabling consistent and repeatable Databricks deployments across all environments. The project encompassed infrastructure automation, data ingestion modernization, and CI/CD pipeline establishment within a modern Lakehouse architecture.
Designed and implemented Terraform modules for the entire Databricks infrastructure, including workspace configuration, cluster policies, job definitions, secret scopes, and Unity Catalog objects — all version-controlled in Git.
Migrated all Hive metastore tables to Unity Catalog with automated schema validation, enabling centralized governance, cross-workspace data sharing, and complete audit trail capabilities.
Deployed Fivetran connectors for automated, reliable data ingestion from Salesforce, replacing fragile custom scripts with a managed service that provides schema drift detection, automatic retry logic, and data freshness monitoring.
Established robust CI/CD pipelines using GitHub Actions integrated with Terraform, enabling infrastructure changes to go through automated plan → review → apply workflows with approval gates.
Built a Lakehouse architecture that organizes data into bronze (raw), silver (cleaned), and gold (business-ready) layers, with automated data quality checks at each transformation stage.
Created self-service Terraform templates that enable AstraZeneca's engineering teams to provision new data pipeline environments in minutes instead of days, with guardrails enforced through policy-as-code.
Implementation Phases
Current State Assessment
3 WeeksAudited all existing Databricks configurations, Hive metastore schemas, custom ingestion scripts, and deployment processes. Documented the gap between current state and target architecture, and prioritized remediation items.
Terraform Foundation
4 WeeksDeveloped core Terraform modules for Databricks workspace provisioning, cluster policy management, job definitions, and Unity Catalog objects. Established the remote state backend, module registry, and CI/CD integration.
Hive to Unity Catalog Migration
4 WeeksMigrated all Hive metastore tables to Unity Catalog in a phased approach, validating data integrity at each step. Implemented automated schema synchronization and access control policies across all environments.
Fivetran Integration & Data Ingestion
3 WeeksDeployed Fivetran connectors for Salesforce and other data sources, configured destination schemas in Unity Catalog, set up freshness monitoring, and decommissioned legacy custom ingestion scripts.
CI/CD Pipeline Implementation
3 WeeksBuilt end-to-end CI/CD pipelines for both infrastructure (Terraform) and data pipeline (Databricks jobs) deployments. Implemented automated testing, approval gates, and rollback mechanisms.
Validation & Knowledge Transfer
3 WeeksPerformed comprehensive end-to-end testing across all environments, documented all Terraform modules and CI/CD processes, and conducted hands-on training sessions for AstraZeneca's engineering team.
Technologies Used
Key Results
Measurable outcomes and business impact delivered through this engagement.
Achieved consistent, repeatable deployments across all Databricks environments (dev, staging, production) with zero configuration drift, reducing deployment-related incidents by 90%.
Eliminated environment drift entirely through Terraform-managed infrastructure, ensuring that every environment is an exact replica of the declared configuration in version control.
Automated data ingestion from Salesforce and other sources via Fivetran, reducing ingestion-related maintenance effort by 80% and improving data freshness from daily to near real-time.
Enhanced data governance capabilities with Unity Catalog's built-in lineage tracking, audit logging, and granular access controls, meeting pharmaceutical regulatory requirements.
Accelerated CI/CD cycles for data pipeline deployment from weeks to hours, with automated testing and approval workflows that maintain quality while increasing velocity.
Enabled self-service environment provisioning, allowing teams to spin up new development environments in under 15 minutes with production-parity configurations.
Before vs. After Comparison
Deployment Time
Before
2-3 weeks
After
2-3 hours
95% faster
Environment Drift
Before
Frequent issues
After
Zero drift
100% eliminated
Data Freshness
Before
Daily batches
After
Near real-time
24x improvement
Ingestion Maintenance
Before
20+ hrs/week
After
4 hrs/week
80% reduction
Env Provisioning
Before
3-5 days
After
15 minutes
99% faster
Deployment Incidents
Before
8-10/month
After
< 1/month
90% reduction
Key Infrastructure Metrics (Post-Implementation)
Effort Distribution by Workstream
Platform Maturity — Before vs After
Infrastructure Automation Coverage (%)
Previous Case Study
Amgen — Enterprise Unity Catalog Migration
Next Case Study
Coca-Cola — Supply Chain Data Pipeline
Facing a Similar Challenge?
Let us help you architect a solution that delivers measurable business impact.
