Google's Spanner Omni Reaches General Availability, Replacing Proprietary Hardware with Software

Distributed SQL database now runs in customer data centers and other clouds, but without an availability SLA

By LineZotpaper
Published
Read Time3 min
Google has released Spanner Omni to general availability, a version of its distributed SQL database that can run in a customer's own data center, on other cloud platforms, or even on a laptop. The release marks a significant engineering effort, replacing the two components most tied to Google's infrastructure: the Colossus distributed file system and the TrueTime clock service, which relied on atomic clocks and GPS. However, the service does not come with an availability service level agreement.

Google has made Spanner Omni generally available, a deploy-anywhere version of its distributed SQL database that runs in a customer's own data center, on other clouds, or on a laptop.

Getting Spanner off Google's infrastructure meant replacing the two components it depended on most: Colossus, the distributed file system, and TrueTime, the clock service built on atomic clocks and GPS. What does not come with it is an availability SLA.

In place of Colossus, Spanner Omni introduces what the company calls a Colossus-like abstraction layer, writing to attached local file systems and making them available across the network to other nodes, with shard splitting and rebalancing handled automatically. Google is direct about the compromise: the file layer is not Colossus, but it is a sufficient stand-in to perform comparably to the managed service for most workloads.

TrueTime got the same treatment. The software-based alternative provides error-bounded time synchronization across servers, as the original does using atomic clocks and GPS. Google's explanation for why that works turns on how Spanner already behaves: the database overlaps time uncertainty waits with other work, so it tolerates weaker uncertainty bounds than TrueTime delivers in practice. That slack is what allows timekeeping across heterogeneous hardware without limiting availability or performance.

Paxos consensus, automatic sharding and synchronous replication carry over unchanged, and Google's internal benchmarks claim millions of queries per second across petabytes in a single regional deployment.

Practitioner reaction has focused on what that means to run. Carlos Pérez Martín, CTO at Q2BSTUDIO, argued that the change is not primarily a design question: "The interesting shift is operational, not architectural: once the same engine runs in your racks, the failure domains become yours."

He set out three consequences. Quorum and witness topology has to be re-derived for the local latency budget, and when a whole datacenter becomes the unit of failure, p99 rather than the mean is what the application feels. Patching, versioned upgrades and rollback stop being someone else's ticket queue and become a change-management problem with the customer's name on it. And moving data residency into a private datacenter brings the backup, key management and audit burden along with it.

His recommended test is narrower than a feature comparison: "A like-for-like pilot against the managed service, measured on tail latency and ops toil instead of feature parity, is the cheap" way to evaluate the offering.

§

Analysis

Why This Matters

  • Organizations that need Spanner's consistency and scalability but cannot run on Google Cloud now have a self-managed option, expanding the database's reach into regulated industries or hybrid architectures.
  • The removal of a cloud SLA means customers bear full operational risk, shifting the cost-benefit calculation significantly compared to the managed service.
  • This approach could set a precedent for other cloud-native databases that have relied on proprietary infrastructure, potentially reshaping enterprise database deployment strategies.

Background

Spanner is Google's globally distributed SQL database, known for providing strong consistency and horizontal scalability across data centers. It originally relied on Google's proprietary Colossus file system and TrueTime, a clock service using atomic clocks and GPS for precise time synchronization. Spanner Omni was first announced as a preview to allow running the database outside Google's cloud, and has now reached general availability. The release represents a technical milestone in decoupling the database from Google's physical infrastructure.

Key Perspectives

Google Cloud: Positions Spanner Omni as a way to bring Spanner's capabilities to any environment, including on-premises data centers, other clouds, and edge deployments, without requiring customers to migrate to Google Cloud.

Enterprise customers and practitioners: See the operational burden as the critical issue. Running a globally distributed database with Paxos consensus and synchronous replication demands careful topology planning, patching, and backup management that was previously handled by Google.

Industry observers: The lack of an availability SLA is a significant differentiator from the managed service, meaning customers accept full responsibility for uptime, which may make the offering suitable for specific use cases but less attractive for mission-critical workloads without dedicated operational expertise.

What to Watch

  • Adoption rates and real-world performance benchmarks comparing Spanner Omni to other self-hosted distributed databases like CockroachDB or YugabyteDB.
  • Google's future roadmap: whether it introduces an SLA or additional managed services for Spanner Omni deployments.
  • Community and vendor ecosystem around enabling operational tooling (backup, monitoring, patching) for Spanner Omni in non-Google environments.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.