Case Study 05Gaming / FinTechAWS

    Enterprise Cost Optimization & Governance for Gaming Platform

    FanDuel, one of America's leading sports gaming and entertainment companies, had seen its Databricks environment grow organically alongside rapid business expansion. Without standardized resource management, monthly Databricks costs had ballooned to $300K with limited visibility into which teams, workloads, or projects were driving spending. Uncontrolled cluster sprawl, inconsistent resource allocation, and inadequate role-based access controls created both financial and governance challenges. Vrahad Analytics was engaged to implement a comprehensive cost optimization and governance strategy that would bring costs under control while maintaining — and even improving — platform performance for all user personas.

    Client

    FanDuel

    Cloud Platform

    AWS

    Duration

    3 Months

    Enterprise Cost Optimization & Governance for Gaming Platform

    The Challenge

    FanDuel's Databricks environment reflected the classic pattern of organic growth outpacing governance. As the company scaled from startup to industry leader, the data platform grew without the guardrails needed to manage costs and access at enterprise scale.

    1

    Monthly Databricks compute costs had reached $300K and were trending upward, with no clear attribution model to understand which teams, projects, or workloads were driving the majority of spending.

    2

    Cluster sprawl was rampant — over 200 active clusters with no standardized sizing, auto-termination policies, or lifecycle management, meaning idle clusters routinely ran 24/7 consuming expensive on-demand compute.

    3

    Inconsistent resource allocation meant that data analysts were running notebooks on the same cluster configurations as heavy ETL workloads, leading to both overspending on analyst clusters and performance issues on shared resources.

    4

    Limited cost visibility across teams made it impossible for engineering managers to make informed decisions about resource allocation or to hold teams accountable for their Databricks consumption.

    5

    Inadequate role-based access controls created security risks — analysts could create arbitrary cluster configurations, and there was no enforcement of organizational standards for compute resources.

    6

    No proactive cost monitoring meant that spending spikes (caused by misconfigured jobs, runaway notebooks, or forgotten clusters) were only discovered at the end of the billing cycle, weeks after the waste occurred.

    Our Solution

    We implemented a comprehensive cost optimization and governance strategy using Terraform-managed cluster policies as the foundation. The approach balanced cost reduction with user experience, ensuring that each persona — analysts, engineers, and data scientists — received appropriately sized resources with automated guardrails that prevent waste without impeding productivity.

    Conducted a thorough cost analysis of the existing environment, profiling every cluster's utilization patterns, idle time, and cost attribution. Identified that 45% of total spend was on idle or significantly underutilized clusters.

    Designed and deployed Terraform-managed cluster policies for three distinct user personas: analysts (small, short-lived interactive clusters), data engineers (medium, auto-scaling batch clusters), and data scientists (GPU-enabled, spot-heavy ML clusters).

    Enforced auto-termination after 1 hour of inactivity across all interactive clusters via cluster policies, eliminating the single largest source of waste — clusters left running overnight and over weekends by users who forgot to shut them down.

    Deployed spot instances for all fault-tolerant workloads (batch ETL, model training, data quality checks), achieving up to 90% compute cost savings on these workloads by leveraging AWS's excess capacity at steep discounts.

    Implemented Unity Catalog for unified governance with least-privilege access controls, ensuring that users could only create clusters that conform to their persona's policy and access only the data assets relevant to their role.

    Built a continuous cost monitoring and alerting system using custom dashboards that provide daily cost breakdowns by team, workload type, and cluster — with automated Slack alerts for spending anomalies that exceed thresholds.

    Implementation Phases

    1

    Cost Discovery & Profiling

    2 Weeks

    Analyzed 90 days of Databricks usage data to profile every cluster's utilization, idle time, and cost. Categorized workloads by type (interactive, batch, ML) and mapped spending to teams and projects. Identified $180K/month in optimization opportunities.

    2

    Cluster Policy Design

    2 Weeks

    Designed three persona-based cluster policies (analyst, engineer, data scientist) with Terraform, defining allowed instance types, auto-scaling ranges, auto-termination settings, spot instance ratios, and tagging requirements for cost attribution.

    3

    Terraform Deployment & Policy Enforcement

    3 Weeks

    Deployed cluster policies via Terraform across all workspaces, migrated existing clusters to policy-compliant configurations, and enabled policy enforcement that prevents creation of non-compliant clusters.

    4

    Spot Instance Strategy & Optimization

    2 Weeks

    Implemented spot instance pools for fault-tolerant workloads with automatic fallback to on-demand for critical jobs. Configured instance fleet strategies to maximize spot availability while minimizing interruption risk.

    5

    Monitoring, Alerting & Governance

    3 Weeks

    Built real-time cost monitoring dashboards, implemented automated Slack alerts for spending anomalies, deployed Unity Catalog for data governance, and conducted training sessions for all user personas on the new policies and self-service capabilities.

    Technologies Used

    TerraformDatabricks Cluster PoliciesUnity CatalogAWS Spot InstancesAutoscalingRBACCost MonitoringAWSPythonSlack API

    Key Results

    Measurable outcomes and business impact delivered through this engagement.

    Reduced monthly Databricks costs from $300K to $120K — a 60% savings ($180K/month, $2.16M annualized) — while maintaining or improving performance for all workload types.

    Implemented role-based cluster policies for 3 distinct user personas (analysts, engineers, data scientists), ensuring each group gets appropriately sized resources without over-provisioning.

    Achieved up to 90% compute cost savings on fault-tolerant workloads through strategic spot instance deployment with automatic fallback to on-demand for business-critical jobs.

    Established continuous cost monitoring with real-time dashboards and automated anomaly alerting, reducing mean time to detect spending spikes from weeks (end of billing cycle) to minutes.

    Enforced least-privilege access and unified governance via Unity Catalog, ensuring that all users operate within organizational guardrails while maintaining self-service capabilities.

    Eliminated cluster sprawl entirely — active clusters reduced from 200+ unmanaged to 80 policy-compliant clusters with automated lifecycle management and utilization tracking.

    Before vs. After Comparison

    Monthly Cost

    Before

    $300,000

    After

    $120,000

    60% reduction

    Annual Savings

    Before

    $3.6M/year

    After

    $1.44M/year

    $2.16M saved

    Active Clusters

    Before

    200+ unmanaged

    After

    80 policy-managed

    60% fewer

    Idle Waste

    Before

    45% of spend

    After

    < 5% of spend

    90% eliminated

    Spot Instance Usage

    Before

    0%

    After

    70% of batch

    Up to 90% savings

    Cost Anomaly Detection

    Before

    End of month

    After

    Real-time

    Minutes vs weeks

    Monthly Databricks Cost ($K)

    BeforeAfterSavings075150225300

    Compute Cost Distribution (After)

    Platform Maturity — Before vs After

    Cost ControlGovernanceVisibilitySelf-ServicePerformanceCompliance0255075100
    Before After

    Monthly Cost Trend ($K)

    Month 1Month 2Month 3Month 4Month 5Month 6075150225300

    Facing a Similar Challenge?

    Let us help you architect a solution that delivers measurable business impact.