Joran from TigerBeetle here!
Awesome to hear that you're excited about Zig.
Thanks for sharing the link to our repo also—would love to take you on a 1-on-1 tour of the codebase sometime if you'd be up for that!
HN user
Joran from TigerBeetle here!
Awesome to hear that you're excited about Zig.
Thanks for sharing the link to our repo also—would love to take you on a 1-on-1 tour of the codebase sometime if you'd be up for that!
Anecdotally again, but I've been coding in Zig since 2020 and have hit I think 2-3 compiler bugs in all that time?
The first was fixed within 24 hours, in fact just before I reported it. The others had clear TODO error messages in the panic, and there were easy enough workarounds.
Yes, we're planning also to add a kill switch to the allocator that we switch on if anything allocates after init().
Ah, missed that, thanks! I've updated the comment.
We're aware of this, in fact, and do have a plan to address virtual memory. To be fair, it's really the kernel being dynamic here, not TigerBeetle.
Static allocation has also made TigerBeetle's code cleaner, by eliminating branching at call sites where before a message might not always have been available. With static allocation, there's no branch because a message is always guaranteed to be available.
It's also made TigerBeetle's code more reliable, because tests can assert that limits are never exceeded. This has detected rare leaks that might otherwise have only been detected in production.
Joran from the TigerBeetle team here.
The limit of 70 lines is actually a slight increase beyond the 60 line limit imposed by NASA's Power of Ten Rules for Safety Critical Software.
In my experience, in every instance where we've refactored an overlong function, the result has almost always been safer.
It's defense-in-depth.
We use what we have available, according to the context: checksums, assertions, hash chains. You can't always use every technique. But anything that can possibly be verified online, we do.
Buffer bleeds also terrify me. In fact, I worked on static analysis tooling to detect zero day buffer bleed exploits in the Zip file format [1].
However, to be clear, the heart of a bleed is a logic error, and therefore even memory safe languages such as JavaScript can be vulnerable.
Sure!
Here's an overview with references to the simulator source, and links to resources from FoundationDB and Dropbox: https://github.com/tigerbeetledb/tigerbeetle/blob/main/docs/...
We also have a $20k bounty that you can take part in, where you can run the simulator yourself.
Thanks!
Joran from the TigerBeetle team here.
Appreciate your balanced comment.
To be fair, we're certainly concerned about logic errors and buffer bleeds. The philosophy in TigerBeetle is always to downgrade a worse bug to a lesser. For example, if it's a choice between correctness and liveness, we'll downgrade the potential correctness bug to a crash.
In the specific case of message buffer reuse here, our last line of defense then is also TigerBeetle's assertions, hash chains and checksums. These exhaustively check all function pre/post-conditions, arguments, processing steps and return values. The assertion-function ratio is then also tracked for coverage, especially in critical sections like our consensus or storage engine.
So—apologies for the wince! I feel it too, this would certainly be a nasty bug if it were to happen.
Joran from the TigerBeetle team here!
We have a secret plan for this too. ;)
Good to be back—Joran from the TigerBeetle team here!
Static allocation does make for some extremely hard guarantees on p100 latency. For example, for a batch of 8191 queries, the performance plot is like Y=10ms, i.e. a flat line.
And memory doesn't increase as throughput increases—it's just another flat line.
I personally find it also to be a fun way of coding, everything is explicit and limits are well-defined.
Thanks!
Data Oriented Design runs like a river through TigerBeetle—it's always on our mind.
By the way, have you seen Andrew Kelley's Handmade Seattle talk on Practical DOD? [1]
Thanks!
Joran from the TigerBeetle team here.
I would imagine this would be potentially more efficient than having conventional mallocs.
Yes, in our experience, static allocation means we use sometimes 10x less memory. For example, TigerBeetle's storage engine can theoretically address up to 100 TiB with less than 1 GiB of RAM for the LSM in-memory manifest, which is efficient.
Because we think about memory so much upfront, in the design phase, we tend not to waste it.
Thanks!
Joran from the TigerBeetle team here.
TigerBeetle's storage engine is designed also for range queries, and we have some interesting ideas for our query engine in the works.
To be sure, there are some tricky things that crop up, such as pipeline blockers for queries. However, we do have limits on literally everything, so that makes it easier—static allocation is best when it's done viral.
We wanted to go into so much more detail on all this but there was only space for so much. Stay tuned.
Joran from the TigerBeetle team here!
TigerBeetle uses Deterministic Simulation Testing to test and keep testing these paths. Fuzzing and static allocation are force multipliers when applied together, because you can now flush out leaks and deadlocks in testing, rather than letting these spillover into production.
Without static allocation, it's a little harder to find leaks in testing, because the limits that would define a leak are not explicit.
Thanks!
Joran from the TigerBeetle team here.
This was in fact one of our motivations for static allocation—thinking about how best to handle overload from the network, while remaining stable. The Google SRE book has a great chapter on this called "Handling Overload" and this had an impact on us. We were thinking, well, how do we get this right for a database?
We also wanted to make explicit what is often implicit, so that the operator has a clear sense of how to provision their API layer around TigerBeetle.
Basically, and we're huge fans of FoundationDB!
As you zoom in, you will see differences. For example, TigerBeetle's data structures are all cache line aligned, and we use static allocation etc. The storage fault model is significantly different.
We use the same testing techniques though. FDB are pioneers in the space.
Thanks, it's a pleasure! Yes, with packed structs it's important to keep things carefully aligned.
Have you seen this post [1] about struct packing? This is what we did for TB's structures, so that we can switch them to `extern struct` for C ABI compatibility. It also helped side step the packed struct bugs.
Thanks @eternalban!
Awesome to read your comment here—appreciate the well wishes!
To be clear, this was out of scope of the bounty, it was a bug in Apple, that TB awarded anyway.
You could one day replace the state machine with your own and have, for example, Redis.
You're right.
And yet there's so much that is there!
The Normal protocol. The View Change protocol. The CTRL protocol from PAR (that you don't get to see often). Thousands of lines of code that are incredibly hard to get right.
All the fault models. The storage fault model alone is also not something you find many distributed systems attempting, let alone paying bounties for.
It's also not common to find bounties that go out of their way to help you. TigerBeetle's bounty ships with a state of the art Deterministic Simulation fuzzing tool that you can use to explore interesting state spaces quicker. It will even classify bugs as liveness or correctness for you. It's like your own Jepsen, except you can inject storage faults, speed up time, and replay anything you find from a seed.
Again, the only reason we were explicit about scope really, is because of our own experience doing bounty programs that were underspecified. For example, while it should be clear enough that this is a distributed systems and consensus bug bounty challenge, literally called “Viewstamped Replication Made Famous”, we didn't want anyone to be confused and think it was a security bug bounty. That's the only reason it's excluded, because we want people to break our consensus. Nevertheless, we do have small awards for interesting findings.
So I hope you'll give it a shot! We'd love to announce and award your findings. For example, why not take on the challenge during HYTRADBOI's database jam?
Thanks!
I always love it when this happens. Watching Phil's reaction!
100%
I've definitely been there! Finding P1s for full read/write access and then seeing the report downgraded to a P3, and having to have the platform arbitrate and bump it back to P1.
However, it was this experience of mine as a part-time security researcher that actually led to us creating the bug bounty program for TigerBeetle's consensus.
For example, if you're looking at another database and find a correctness bug, there might not be a bounty program at all. Whereas with TigerBeetle, there hasn't been a single valid report that we haven't awarded, at least so far.
It's also why we were careful to rather be upfront and explicit about scope, than disappoint anyone after the fact.
And we recognize that consensus is hard and takes time to learn, hence the $8192 award for correctness finds.
That said, I hope you can see from the leaderboard that we've been generous. For example, Alex Miller found a bug in Apple's O_DSYNC and we nevertheless awarded $1024 because it was such a great find (Apple thought so too!).
See also: https://news.ycombinator.com/item?id=32788840
TigerBeetle is designed to keep running even if all machines are experiencing radioactive levels of local storage corruption, or else shutdown safely when it detects that it must. We use automated testing to test TigerBeetle with read/write storage fault injection levels as high as 20-30%. On the other hand, this invariant is not typically given by other engines, per the storage fault research that's come out of UW-Madison the past few years. For example, “Protocol-Aware Recovery for Consensus-Based Storage”.
While other engines may have incredible test suites built up over years and years, they were also designed mostly before the advent of autonomous Deterministic Simulation Testing (think Jepsen except you can speed up time and replay bugs, for example, that would otherwise take 10 years to manifest in real time), which is a showstopper.
Finally, we wanted TigerBeetle to be highly available and distributed. TigerBeetle can run across 3 availability zones with 2 replicas in each, with seamless failover. You can stay running even if you lose a whole AZ plus another replica, thanks to Heidi Howard's Flexible Quorums. Again, this problem is not as simple as simply slapping on RAFT for distribution, because “Redundancy Does Not Imply Fault-Tolerance”, and because RAFT makes concessions around storage faults and dueling leaders for the sake of the readability of the paper, that we didn't want to make for TigerBeetle's actual implementation, hence our choice of Viewstamped Replication (MIT, '88, '12).
We think carefully about this.
We do in fact have some extremely advanced testing infrastructure, even going as far as using a deterministic Linux hypervisor to do coverage guided fuzzing of our compiled binaries from the outside in. At the same time, we fuzz from the inside out, also Deterministic Simulation Testing, with a ton of assertions (literally a 1000+ and counting) as a force multiplier for fuzzing. Our experience in all this, is that while Zig is early in terms of timeline, the quality is nevertheless extremely high.
Andrew and team know what they're doing. They've got some of the best people in the world in their respective fields. For example, Frank Denis of libsodium heading up Zig's crypto.
Beyond this, we do the basic things like consciously restrict our use of language features to stable features only (which may be why our experience has not been the same as yours with respect to compiler bugs?), and invest in our own I/O stack around io_uring instead of depending on the std lib, which we know will churn.
Two of our team are Zig core team members and we sponsor the Zig Software Foundation to invest in the ecosystem. Don't forget also that Zig's ecosystem is really C's ecosystem so there's an escape hatch there.
Also, TB is not yet production ready. We'll ship our production release when we're confident that our TB binary is safe, or able to shut down if it detects any safety violation.
At the end of the day, databases are a big investment. It's important to think about the next 20 to 30 years. C would absolutely have been the wrong choice for that future.
As you say, the future is bright for Zig, and we're convinced that Zig is the right choice for a database that follows embedded coding standards with respect to memory. For example, TigerBeetle has to handle memory allocation failure, and only does static memory allocation so we never malloc or free after startup.