Gergely Orosz @gergely.pragmaticengineer.com · Jun 11

The postmortem from Coinbase's 10-hour outage is out and... damn They run global trading from a single region because of latency. OK, I understand. BUT they have no automated failover prepared! Are they praying the region never goes down?? Doesn't compute for me...

60 likes 10 replies

?

Replies

Justin Garrison · Jun 11

Doesn’t compute for them either

N

@nachman16.bsky.social · Jun 11

They probably estimated the cost of paying for zone fail over exceeds the cost of outages they might have

Leo Lapworth · Jun 11

@gergely.pragmaticengineer.com Ummm, that screen shot says `zone` - not `region` - it's worse.. They might still do it for latency, but zones in a region are geographically near. AWS does lots of automatic failover within region, they have specifically engeneered it not to do that.

Redowan Delowar · Jun 11

MR is fiendishly hard to pull off. Every discussion w/ MR only talks about the DB and calls it a day. While proper DB failover is hard enough, failing over brokers, queues are harder. Even teams that pull it off, w/o proper chaos testing those code path never get visited until a region goes down.

CTO · Jun 11

Not even region, the ability to failover between AZs..

Sagittarius A* · Jun 11

> serious market > coinbase

Joseph Gruber · Jun 11

Not even single region but single AZ it seems

Jeremy Lewi · Jun 11

I thought the entire point of crypto was that it was decentralized and permission less.

Gergely Orosz · Jun 11

Full postmortem: www.coinbase.com/en-gb/blog/a... I would get this for a small company. But a $40B company? It's a serious WTH moment for me. And my impression of Coinbase engineering nosedived. This is resilliency basics: if you depend on a region, have auto failover and exercise it...

Cryptohopper · Jun 11

Single-region architecture is risky for any trading platform. Automated systems need redundancy baked in at the protocol level, not just failover plans.