DataCenterNews Canada - Specialist news for cloud & data center decision-makers
Canada
Google flexible VMs aim to keep Spark jobs running

Google flexible VMs aim to keep Spark jobs running

Tue, 6th Oct 2026 (Today)
Mara Sugue
MARA SUGUE News Editor

Google has outlined the use of flexible virtual machines in its Managed Service for Apache Spark to reduce cluster provisioning failures during compute shortages.

Rising demand for AI infrastructure has increased pressure on compute capacity, creating problems for Spark users when a preferred machine family is unavailable in a given zone or region. In those cases, rigid infrastructure choices can delay cluster creation, interrupt job execution and affect service targets for time-sensitive analytics pipelines.

Flexible VMs let users define an ordered list of acceptable machine families for master, primary worker and secondary worker nodes, rather than tying a cluster to a single instance type. The feature allows a mix of machine types and generations in one configuration, including Gen2 families such as N2 and N2D alongside Gen4 families including N4 and C4.

It also supports storage choices that adapt to the machine family selected at provisioning. These rules can be applied across the cluster, including primary workers, spot or preemptible secondary workers, and master nodes.

Ranked setup

A central part of the model is a ranking system that determines the order in which machine families are tried. Teams should specify at least two machine families in the top-priority rank to improve the chances of obtaining capacity when demand is high.

For production pipelines built around n2d-standard-16 shapes, Google set out a tiered example that starts with n2d-standard-16 and n2-standard-16 at Rank 0, then moves to n4-standard-16 and n4d-standard-16 at Rank 1. It then falls back to c4-standard-16 and c3-standard-22 at Rank 2, and e2-standard-16 at Rank 3.

A separate example for workloads using legacy n1-standard-16 shapes places n1-standard-16 and n2-standard-16 in the first rank, followed by n2d-standard-16, then n4-standard-16 and n4d-standard-16, and finally e2-standard-16. This gives users a path to newer machine architectures while maintaining operational continuity.

Storage choices

The guidance places particular emphasis on Hyperdisk Balanced for newer instance families. N4 and C4 machines depend on Hyperdisk to provide predictable performance across varying VM sizes, and default IOPS and throughput settings are likely to suit many distributed Spark jobs.

The broader point is that storage and compute choices now need to be planned together. If users want broader fallback options, they may also need to accept a change in disk type when workloads move between older and newer machine families.

Operational trade-offs

Google also highlighted several issues companies must consider when broadening their machine options. One is quotas. Instead of maintaining quota for a single machine family, organisations need enough compute and disk quota for every machine type and storage option listed in their flexible VM rankings, including Hyperdisk.

Committed spending is another issue. Traditional committed use discounts are usually tied to specific machine families, which limits the value of switching between architectures during shortages. Users seeking savings across multiple VM families and regions should consider compute flexible committed use discounts instead.

Performance also becomes less predictable across mixed infrastructure. Results can differ between machine generations and between storage options such as local SSD and Hyperdisk, so users should test their own Spark jobs to assess the effect on service levels.

Broader resilience steps

Alongside flexible VMs, Google recommended a series of operational measures to improve the chances of jobs running under constrained capacity conditions. These include using automatic zone selection so the service can place jobs in a zone with available resources, favouring smaller machine shapes rather than heavily contested large-core instances, and using autoscaling with sensible upper limits for variable workloads.

Google also pointed to partial cluster creation, which allows a cluster to start with a minimum acceptable number of primary workers and add more later as capacity becomes available. Regional fallback planning is another part of the guidance, particularly for high-demand areas such as us-central1.

The push reflects a wider challenge for cloud users as AI workloads compete with conventional analytics and data processing for the same underlying infrastructure. For Spark users running scheduled or business-critical pipelines, the question is no longer only which instance type best suits a workload, but how many acceptable substitutes can be built into the configuration before a shortage leads to failure.

Flexible VMs are intended to help keep Spark pipelines running when regional or zonal stockouts affect a preferred machine family.