Experimenting with celld in increasingly pathological situations. After merging the replicated write-behind log, this experiment had 4 instances of celld, each on its own machine, and round-robin kills them every few minutes to see how the fleet handles it.
Testing is done at a few layers: a TLA+ model, deterministic simulation testing, typical regression testing, and real life lab experiments like this.
Maybe some day I'll be able to experiment with it at discord scale :)