Case Study 01PharmaceuticalsAWS

    Enterprise Databricks Implementation & Unity Catalog Migration

    Amgen, one of the world's leading biotechnology companies, needed to modernize its data infrastructure from a fragmented AWS Glue environment to a centralized, governed Databricks ecosystem. With over 400 data pipelines distributed across multiple databases and teams, the migration required surgical precision to avoid disrupting ongoing drug development and clinical operations. Vrahad Analytics was engaged to plan, execute, and validate this large-scale transformation, delivering a unified data platform that set the foundation for Amgen's next-generation analytics capabilities.

    Client

    Amgen Pharma Inc.

    Cloud Platform

    AWS

    Duration

    6 Months

    Enterprise Databricks Implementation & Unity Catalog Migration

    The Challenge

    Amgen's data infrastructure had grown organically over several years, resulting in a sprawling AWS Glue environment that was becoming increasingly difficult to manage, govern, and scale. The pharmaceutical industry's stringent regulatory requirements made the situation even more complex.

    1

    Over 400 data pipelines were distributed across multiple AWS Glue databases with no centralized catalog, making it nearly impossible to discover, audit, or govern data assets at scale.

    2

    Fragmented metadata management meant that different teams maintained their own schema definitions, column naming conventions, and data quality standards, creating widespread inconsistency.

    3

    Inconsistent access controls across databases posed significant compliance risks for a pharmaceutical company subject to FDA and HIPAA regulations, with no unified permission model.

    4

    The lack of data lineage visibility made it extremely difficult to perform impact analysis when upstream schemas changed, often causing downstream pipeline failures.

    5

    External tables and mount points created complex dependencies between S3 storage and Glue catalogs, requiring careful handling during any migration to avoid breaking references.

    6

    Business operations could not tolerate any downtime — ongoing drug research, clinical trial data processing, and commercial analytics all depended on continuous pipeline execution.

    Our Solution

    We orchestrated a full-scale migration from AWS Glue to Databricks Unity Catalog using a carefully designed phased approach. Every step was validated against Amgen's operational requirements, with dual administration maintained throughout the transition to ensure business continuity.

    Conducted a comprehensive assessment of the entire AWS Glue environment, cataloging all 400+ pipelines along with their dependencies, schedules, data sources, and downstream consumers to build a complete migration map.

    Developed a strategic Unity Catalog adoption plan that organized data assets into a three-level namespace (catalog → schema → table) aligned with Amgen's organizational structure and access control requirements.

    Implemented careful handling of external tables and mount points, creating a migration path that preserved existing S3 data references while transitioning metadata ownership to Unity Catalog.

    Established dual administration protocols across Databricks and AWS, enabling both platforms to operate simultaneously during the migration window with synchronized access policies.

    Built automated validation scripts that compared source (Glue) and target (Unity Catalog) metadata, row counts, checksums, and schema definitions for every migrated table and pipeline.

    Designed a rollback strategy for each migration batch, ensuring that any issues detected during validation could be quickly reversed without impacting production workloads.

    Implementation Phases

    1

    Discovery & Assessment

    4 Weeks

    Cataloged all 400+ AWS Glue pipelines, mapped dependencies, identified external tables and mount points, and assessed data quality baselines. Produced a comprehensive migration readiness report with risk ratings for each pipeline.

    2

    Architecture Design & Planning

    3 Weeks

    Designed the Unity Catalog namespace hierarchy, defined access control policies aligned with Amgen's regulatory requirements, planned the migration sequence (prioritizing low-risk pipelines first), and established success criteria.

    3

    Infrastructure Setup & Pilot

    3 Weeks

    Provisioned the Databricks workspace with Unity Catalog enabled, configured AWS IAM roles for cross-service access, set up the dual administration framework, and migrated the first batch of 50 pipelines as a pilot to validate the approach.

    4

    Batch Migration Execution

    10 Weeks

    Executed the full migration in batches of 40–60 pipelines per week. Each batch followed a migrate → validate → monitor → approve cycle. Automated validation scripts ran against every table to ensure data integrity.

    5

    Validation & Optimization

    3 Weeks

    Performed end-to-end validation of all migrated pipelines, optimized data access patterns, fine-tuned cluster configurations for different workload types, and verified governance policies were correctly enforced across all catalogs.

    6

    Cutover & Decommissioning

    2 Weeks

    Executed final cutover from AWS Glue to Databricks as the primary orchestration platform, decommissioned redundant Glue jobs, and transferred operational ownership to Amgen's internal data engineering team with full documentation and runbooks.

    Technologies Used

    AWS GlueDatabricks Unity CatalogAWS S3Delta LakePySparkAWS IAMDatabricks WorkflowsSQL WarehousesPythonAWS CloudFormation

    Key Results

    Measurable outcomes and business impact delivered through this engagement.

    Successfully migrated all 400+ data pipelines from AWS Glue to Databricks with zero disruption to ongoing business operations, drug research, or clinical trial data processing.

    Implemented Unity Catalog as the single source of truth for data governance, providing centralized access controls, audit logging, and data lineage across the entire organization.

    Optimized data access patterns across AWS and Databricks ecosystems, reducing average query execution time by 35% through better partitioning, caching, and Delta Lake optimization.

    Reduced metadata management overhead by 70% through the unified catalog, eliminating duplicate schema definitions and standardizing naming conventions across all teams.

    Established a scalable governance framework that supports Amgen's future data growth, with automated policy enforcement for new tables, schemas, and catalogs.

    Delivered comprehensive documentation and runbooks, enabling Amgen's team to independently manage and extend the platform post-engagement.

    Before vs. After Comparison

    Pipelines Migrated

    Before

    0 on Databricks

    After

    400+ on Unity Catalog

    100% migration

    Data Governance

    Before

    Fragmented / Manual

    After

    Centralized / Automated

    Unified governance

    Query Performance

    Before

    Baseline

    After

    35% faster

    35% improvement

    Metadata Overhead

    Before

    High (manual)

    After

    Minimal (automated)

    70% reduction

    Business Disruption

    Before

    N/A

    After

    None

    Zero downtime

    Compliance Posture

    Before

    Inconsistent

    After

    Fully auditable

    100% coverage

    Migration Scale — Assets Migrated

    PipelinesTablesSchemasUsers03006009001200

    Pipeline Distribution by Type

    Platform Capability — Before vs After

    GovernancePerformanceScalabilitySecurityLineageAutomation0255075100
    Before After

    Cumulative Pipelines Migrated Over Time

    Week 1Week 4Week 8Week 12Week 16Week 240150300450600

    Facing a Similar Challenge?

    Let us help you architect a solution that delivers measurable business impact.