Unexpected Outages and Bug Identification
Users of the peer-to-peer networking service Tailscale may have experienced some unexpected outages starting late last year. A thorough six-month investigation ultimately pinpointed the source: A bug in SQLite’s write-ahead log that had remained concealed for 16 years.
SQLite Bug Revealed
The Tailscale team shared in a Wednesday blog entry that they had finally resolved the issue with assistance from SQLite maintainers, who even needed to develop a new tool (financed by Tailscale) to monitor virtual file system activity in order to locate the problem. Tailscale software engineer Alex Chan characterized it as evading “all our initial efforts to detect it.”
Comprehending the Tailscale Platform
To grasp what transpired, it’s essential to understand how Tailscale operates. The service, utilizing the WireGuard VPN protocol, directly links devices within a virtual private mesh network. It aims to maintain low complexity and ease of implementation for various tasks, from remotely accessing a NAS to linking teams within a unified private network.
The Function of SQLite
Every mesh network, or “tailnet,” resides on one of several servers, where a SQLite database oversees all information regarding the tailnets it contains. “We’ve employed SQLite as our primary database since 2022, and we selected it for its reputation, reliability, and widespread use,” Chan noted in the company’s post-mortem.
Backups and Database Failure
However, a year ago, something went drastically awry. “In our present backup pipeline, we capture a complete snapshot of the database every few minutes, then transmit the entire SQLite file to an S3 bucket,” Chan explained. By August 2025, those backups started encountering database corruption consistently, with no discernible common cause.
Unsuccessful Reproduction Efforts
The Tailscale team was unable to replicate the issue due to the absence of reliable triggers. No low-level code had been altered for months. A review of all interactions with SQLite yielded no results.
The WAL-Reset Bug Surfaces
Suspicions began to focus on SQLite’s checkpointing mechanism. SQLite offers an option to enhance performance and concurrency known as the Write-Ahead Log (WAL), which serves as the previously mentioned hopper.
Database Checkpoint Conflict
Characterized by Chan as “a rare data race in the SQLite source code between a checkpoint and write transaction,” it essentially represents a conflict between checkpoints and writing data to the WAL.
Unlikely Event
According to the SQLite team’s WAL-Reset documentation, the problem can only be activated when WAL mode is enabled and multiple database connections are open on the same file.
SQLite Bug’s Background and Fix
SQLite maintainers suspect the bug dates back to version 3.7.0, which was released in July 2010. It has now been resolved, and the SQLite team advises users to upgrade to a fixed version.
Insights Gained
This incident serves as a valuable reminder for developers: Even the most mundane, dependable software can pose risks when utilized in unusual ways. “Most users operate SQLite in a standard setup and never encounter such issues,” Chan remarked.
Conclusion
SQLite Bug: Concealed Longer Than a Geordie from the Rain!