ЁЯПл The SchoolтА║ЁЯМР Distributed SystemsтА║ЁЯУЪ рдзрдбрд╛ 05 тАФ Replication: рдиреЛрдВрджрд╡рд╣реАрдЪреНрдпрд╛ рдкреНрд░рддреА
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУЪ рдзрдбрд╛ 05 тАФ Replication: рдиреЛрдВрджрд╡рд╣реАрдЪреНрдпрд╛ рдкреНрд░рддреА

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 05 ┬╖ рдорд╛рдЧреЗ: lesson-04-failure-detection ┬╖ рдкреБрдвреЗ: lesson-06-quorums


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ04, рдЕрдзрд┐рдХ рднрд╛рдЧ 2 рд╕реБрд░реВ рд╣реЛрддреЛ: рдкреНрд░рддреА. рдПрдХ рд╢рд╛рдЦрд╛ leader рдЕрд╕рддреЗ рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ write рдШреЗрддреЗ; рдмрд╛рдХреАрдЪреНрдпрд╛ followers рдЕрд╕рддрд╛рдд рдЖрдгрд┐ рддрд┐рдЪрд╛ log copy рдХрд░рддрд╛рдд. Synchronous, asynchronous рдЖрдгрд┐ semi-synchronous replication, replication lag, failover рдордзреНрдпреЗ рдХрд╛рдп рд╣рд░рд╡рддреЗ, рдЖрдгрд┐ split brain. dist/demo.py рдордзрд▓реЗ replication() рдЖрдгрд┐ dist/sim.py рдордзрд▓реЗ Leader.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдкреБрдгреЗ рдореБрдЦреНрдп рдиреЛрдВрджрд╡рд╣реА ЁЯУТ рдареЗрд╡рддреЗ. рдлрдХреНрдд рдкреБрдгреЗрдЪ рддреНрдпрд╛рдд рдЧреБрдг рд▓рд┐рд╣рд┐рддреЗ. рдирд╛рд╢рд┐рдХ рдЖрдгрд┐ рдирд╛рдЧрдкреВрд░ рдкреНрд░рддреА ЁЯУЪ рдареЗрд╡рддрд╛рдд, рдореНрд╣рдгрдЬреЗ рдкреБрдгреНрдпрд╛рдд рдкреВрд░ рдЖрд▓рд╛ рддрд░реА рд╕рдЧрд│реЗ рдЧреБрдг рдирд╖реНрдЯ рд╣реЛрдд рдирд╛рд╣реАрдд.

рдкреБрдгреНрдпрд╛рддрд▓реА рдХрддрд░рд┐рдирд╛ рдЪрд╛рд░ рдЧреБрдг рд▓рд┐рд╣рд┐рддреЗ: g1, g2, g3, g4. рдкреНрд░рддреАрдВрдирд╛ рддреЗ рдХрд╕реЗ рдХрд│рддрд╛рдд?

рдЖрдгрд┐ рдЖрдгрдЦреА рдПрдХ рдзреЛрдХрд╛: рдкреБрд░рд╛рдирдВрддрд░ рдкреБрдгреЗ рдХреЛрд░рдбреЗ рд╣реЛрддреЗ рдЖрдгрд┐ рдЕрдЬреВрдирд╣реА рд╕реНрд╡рддрдГрд▓рд╛ рдореБрдЦреНрдп рдиреЛрдВрджрд╡рд╣реА рд╕рдордЬрддреЗ. рдЖрддрд╛ рджреЛрди рд╢рд╛рдЦрд╛ writes рдШреЗрддрд╛рдд. рдпрд╛рд▓рд╛рдЪ split brain рдореНрд╣рдгрддрд╛рдд.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    t["ЁЯСйтАНЁЯПл teacher"] -->|"write g1 g2 g3 g4"| lead["ЁЯУТ Pune тАФ leader<br/>says 'acknowledged' at once"]
    lead -->|"shipped: g1 g2"| f1["ЁЯУЪ Nashik тАФ follower"]
    lead -->|"shipped: g1 g2"| f2["ЁЯУЪ Nagpur тАФ follower"]
    lead --> x["ЁЯТе Pune dies"]
    x --> nl["new leader has g1 g2<br/>lost: g3 g4 тАФ already acknowledged"]
    s["synchronous: 'acknowledged' only after a follower has it<br/>lost: nothing"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/distributed-systems/lesson-diagrams.html#l05

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг "database рд▓рд╛ replica рдЖрд╣реЗ" рдпрд╛рдЪрд╛ рдЕрд░реНрде рдЕрдиреЗрдХрджрд╛ "рдЖрдкрдг data рдЧрдорд╛рд╡реВ рд╢рдХрдд рдирд╛рд╣реА" рдЕрд╕рд╛ рдШреЗрддрд▓рд╛ рдЬрд╛рддреЛ тАФ рдЖрдгрд┐ asynchronous replication рд╡ automatic failover рдЕрд╕рддрд╛рдирд╛ рддреЗ рдЦреЛрдЯреЗ рдЖрд╣реЗ. 2012 рдЪреА рдПрдХ рдкреНрд░рд╕рд┐рджреНрдз GitHub рдШрдЯрдирд╛: рдЬреБрдирд╛ рдЭрд╛рд▓реЗрд▓рд╛ рдПрдХ MySQL follower promote рдХреЗрд▓рд╛ рдЧреЗрд▓рд╛; рддреНрдпрд╛рдиреЗ рдЬреБрдиреНрдпрд╛ leader рдиреЗ рдЖрдзреАрдЪ рджрд┐рд▓реЗрд▓реНрдпрд╛ primary keys рдкреБрдиреНрд╣рд╛ рд╡рд╛рдкрд░рд▓реНрдпрд╛, рдЖрдгрд┐ рддреНрдпрд╛рдЪ keys рдПрдХрд╛ Redis store рдордзреНрдпреЗрд╣реА рд╡рд╛рдкрд░рд▓реНрдпрд╛ рд╣реЛрддреНрдпрд╛, рддреНрдпрд╛рдореБрд│реЗ рдХрд╛рд╣реА рдЦрд╛рд╕рдЧреА data рдЪреБрдХреАрдЪреНрдпрд╛ users рдирд╛ рджрд┐рд╕рд▓рд╛. Sync рдЖрдгрд┐ async рдордзрд▓реА рдирд┐рд╡рдб рдореНрд╣рдгрдЬреЗ write latency рдЖрдгрд┐ acknowledged writes рдЧрдорд╛рд╡рдгреЗ рдпрд╛рдВрдЪреНрдпрд╛рддрд▓реА рдирд┐рд╡рдб, рдЖрдгрд┐ рддреА рддреБрдореНрд╣реА рдЬрд╛рдгреАрд╡рдкреВрд░реНрд╡рдХ рдХрд░рд╛рдпрд▓рд╛ рд╣рд╡реА.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

dist/sim.py рдордзрд▓реЗ Leader(followers, sync=False) рдПрдХ log рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ follower рд╕рд╛рдареА рдПрдХ list рдареЗрд╡рддреЗ. write(item) log рдордзреНрдпреЗ рдЬреЛрдбрддреЗ рдЖрдгрд┐ рд▓рдЧреЗрдЪ 'acknowledged' рдкрд░рдд рджреЗрддреЗ; sync=True рдЕрд╕реЗрд▓ рддрд░ рддреЗ рдЖрдзреА рдкреНрд░рддреНрдпреЗрдХ follower рдордзреНрдпреЗрд╣реА рдЬреЛрдбрддреЗ. ship(upto) log рдордзрд▓реНрдпрд╛ рдкрд╣рд┐рд▓реНрдпрд╛ upto entries followers рдХрдбреЗ copy рдХрд░рддреЗ (async рдирд┐рд░реЛрдкреНрдпрд╛). fail_over() рд╕рд░реНрд╡рд╛рдд рд▓рд╛рдВрдм рдкреНрд░рдд рдЕрд╕рд▓реЗрд▓реНрдпрд╛ follower рд▓рд╛ promote рдХрд░рддреЗ рдЖрдгрд┐ (its copy, the writes it lacks) рдкрд░рдд рджреЗрддреЗ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 dist/demo.py replication
python3 - <<'EOF'
import sys; sys.path.insert(0, "dist"); from sim import Leader
for shipped in (0, 2, 3, 4):
    L = Leader(["nashik", "nagpur"])
    for g in ("g1", "g2", "g3", "g4"): L.write(g)
    L.ship(shipped)
    survivor, lost = L.fail_over()
    print(f"async, {shipped} of 4 shipped before the crash тЖТ new leader has {survivor}, lost {lost}")
EOF

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

replication рд╣реЗ print рдХрд░рддреЗ:

тФАтФА asynchronous replication: 4 grades acknowledged, then Pune (leader) dies тЖТ new leader has ['g1', 'g2'], lost ['g3', 'g4']
тФАтФА synchronous replication: 4 grades acknowledged, then Pune (leader) dies тЖТ new leader has ['g1', 'g2', 'g3', 'g4'], lost []

рддреБрдордЪрд╛ snippet рд╣реЗ print рдХрд░рддреЛ:

async, 0 of 4 shipped before the crash тЖТ new leader has [], lost ['g1', 'g2', 'g3', 'g4']
async, 2 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2'], lost ['g3', 'g4']
async, 3 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2', 'g3'], lost ['g4']
async, 4 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2', 'g3', 'g4'], lost []

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

Async replication рдордзреНрдпреЗ, failover рдордзреНрдпреЗ рдЬреЗ рд╣рд░рд╡рддреЗ рддреЗ рдореНрд╣рдгрдЬреЗ crash рдЪреНрдпрд╛ рдХреНрд╖рдгреАрдЪрд╛ рдиреЗрдордХрд╛ replication lag: 0 рдкрд╛рдард╡рд▓реЗ тЖТ рдЪрд╛рд░рд╣реА acknowledged рдЧреБрдг рд╣рд░рд╡рд▓реЗ; 4 рдкрд╛рдард╡рд▓реЗ тЖТ рдПрдХрд╣реА рдирд╛рд╣реА. рд╢рд┐рдХреНрд╖рд┐рдХреЗрдиреЗ рдЪрд╛рд░рд╣реА рд╡реЗрд│рд╛ "acknowledged" рдРрдХрд▓реЗ рд╣реЛрддреЗ. Synchronous replication рдордзреНрдпреЗ рдХрд╛рд╣реАрдЪ рд╣рд░рд╡рд▓реЗ рдирд╛рд╣реА, рдХрд╛рд░рдг "acknowledged" рдиреЗ рдПрдХрд╛ рдкреНрд░рддреАрдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рд┐рд▓реА рд╣реЛрддреА.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

PostgreSQL тАФ streaming replication default рдиреБрд╕рд╛рд░ asynchronous рдЕрд╕рддреЗ. On a real account, primary рд╡рд░рдЪреНрдпрд╛ postgresql.conf рдордзреНрдпреЗ рджреЛрди standbys рдкреИрдХреА рдПрдХ synchronous рдХрд░рд╛:

synchronous_standby_names = 'FIRST 1 (nashik, nagpur)'   # wait for the first one in the list that is connected
synchronous_commit = on                                  # a commit waits until that standby has flushed the WAL to disk

рдХрд┐рдВрд╡рд╛ 'ANY 1 (nashik, nagpur)' тАФ рдЬреЛ рдЖрдзреА рдЙрддреНрддрд░ рджреЗрдИрд▓ рддреНрдпрд╛рдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рд╛ (quorum commit). synchronous_commit рд╣реЗ remote_apply (standby рдиреЗ рддреЗ рд▓рд╛рдЧреВ рдХрд░реЗрдкрд░реНрдпрдВрдд рдерд╛рдВрдмрд╛, рдореНрд╣рдгрдЬреЗ рддрд┐рдерд▓реЗ reads рддреЗ рдкрд╛рд╣рддрд╛рдд), remote_write, local рдЖрдгрд┐ off рд╕реБрджреНрдзрд╛ рд╕реНрд╡реАрдХрд╛рд░рддреЗ, рдЖрдгрд┐ рддреЗ рдкреНрд░рддреНрдпреЗрдХ transaction рд╕рд╛рдареА рд╡реЗрдЧрд│реЗ рдард░рд╡рддрд╛ рдпреЗрддреЗ тАФ рдЙрджрд╛рд╣рд░рдгрд╛рд░реНрде рдЧрдорд╛рд╡рд▓реА рддрд░реА рдЪрд╛рд▓реЗрд▓ рдЕрд╢рд╛ log line рд╕рд╛рдареА SET LOCAL synchronous_commit = off;. Lag рд╡рд░ рд▓рдХреНрд╖ рдареЗрд╡рд╛:

SELECT application_name, state, sync_state, write_lag, flush_lag, replay_lag
FROM pg_stat_replication;

MySQL тАФ semisynchronous replication рд╣реЗ рдПрдХ plugin рдЖрд╣реЗ: рдХрд┐рдорд╛рди рдПрдХрд╛ replica рд▓рд╛ transaction рдорд┐рд│реЗрдкрд░реНрдпрдВрдд source рдерд╛рдВрдмрддреЛ (rpl_semi_sync_source_wait_for_replica_count).

Managed failover тАФ Amazon RDS Multi-AZ рджреБрд╕рд▒реНрдпрд╛ Availability Zone рдордзреНрдпреЗ рдПрдХ synchronous standby рдареЗрд╡рддреЗ рдЖрдгрд┐ рддреНрдпрд╛рд╡рд░ fail over рдХрд░рддреЗ; RDS read replicas asynchronous рдЕрд╕рддрд╛рдд рдЖрдгрд┐ рдЖрдкреЛрдЖрдк fail over рдХрд░рдд рдирд╛рд╣реАрдд. Patroni рд╕рд╛рд░рдЦреА tools рдПрдХрдЪ leader рдХреЛрдг рд╣реЗ рдард░рд╡рдгреНрдпрд╛рд╕рд╛рдареА etcd рдХрд┐рдВрд╡рд╛ ZooKeeper (рдзрдбреЗ 09тАУ10) рд╡рд╛рдкрд░реВрди PostgreSQL failover рдЪрд╛рд▓рд╡рддрд╛рдд.

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдкреНрд░рддреНрдпреЗрдХ database рд╕рд╛рдареА рд▓рд┐рд╣реВрди рдареЗрд╡рд╛: sync рдХреА async, рдХрд┐рддреА synchronous рдкреНрд░рддреА, failover рдХреЛрдг рдард░рд╡рддреЛ, рдЖрдгрд┐ рдЬреБрдиреНрдпрд╛ leader рд▓рд╛ writes рдШреЗрдгреНрдпрд╛рдкрд╛рд╕реВрди рдХрд╕реЗ рд░реЛрдЦрд▓реЗ рдЬрд╛рддреЗ. рдордЧ failover drill рдЪрд╛рд▓рд╡рд╛ рдЖрдгрд┐ рдХрд╛рдп рд╣рд░рд╡рд▓реЗ рддреЗ рдореЛрдЬрд╛.

тПня╕П рдкреБрдвреЗ

рдПрдХрдЪ leader рдореНрд╣рдгрдЬреЗ bottleneck рдЖрдгрд┐ рдирд┐рд░реНрдгрдпрд╛рдЪреЗ рдПрдХрдЪ рдард┐рдХрд╛рдг. рджреБрд╕рд░рд╛ рдорд╛рд░реНрдЧ: рдЕрдиреЗрдХ рдкреНрд░рддреАрдВрд╡рд░ рд▓рд┐рд╣рд╛, рдЕрдиреЗрдХ рдкреНрд░рддреАрдВрдХрдбреВрди рд╡рд╛рдЪрд╛, рдЖрдгрд┐ рдЖрдХрдбреНрдпрд╛рдВрдирд╛рдЪ рд╣рдореА рджреЗрдК рджреНрдпрд╛ рдХреА рддреБрдореНрд╣рд╛рд▓рд╛ рд╕рд░реНрд╡рд╛рдд рдирд╡реЗ рджрд┐рд╕реЗрд▓. Quorums.

git checkout lesson-06-quorums

ЁЯУЪ Lesson 05 тАФ Replication: copies of the register

ЁЯУН You are here: Lesson 05 of 12 ┬╖ Previous: lesson-04-failure-detection ┬╖ Next: lesson-06-quorums


ЁЯУж What's in this branch

Lessons 01тАУ04, plus Part 2 begins: copies. One branch is the leader and takes every write; the others are followers and copy its log. Synchronous, asynchronous and semi-synchronous replication, replication lag, what a failover loses, and split brain. replication() in dist/demo.py and Leader in dist/sim.py.

ЁЯзТ Explain like I'm 5

Pune keeps the main register ЁЯУТ. Only Pune writes grades in it. Nashik and Nagpur keep copies ЁЯУЪ, so a flood in Pune does not destroy every grade.

Katrina in Pune writes four grades: g1, g2, g3, g4. How do the copies learn about them?

And one more danger: after the flood, Pune dries out and still thinks it is the main register. Now two branches take writes. That is split brain.

ЁЯЧ║я╕П Diagram

flowchart LR
    t["ЁЯСйтАНЁЯПл teacher"] -->|"write g1 g2 g3 g4"| lead["ЁЯУТ Pune тАФ leader<br/>says 'acknowledged' at once"]
    lead -->|"shipped: g1 g2"| f1["ЁЯУЪ Nashik тАФ follower"]
    lead -->|"shipped: g1 g2"| f2["ЁЯУЪ Nagpur тАФ follower"]
    lead --> x["ЁЯТе Pune dies"]
    x --> nl["new leader has g1 g2<br/>lost: g3 g4 тАФ already acknowledged"]
    s["synchronous: 'acknowledged' only after a follower has it<br/>lost: nothing"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/distributed-systems/lesson-diagrams.html#l05

тЭУ What

ЁЯдФ Why

Because "the database has a replica" is often read as "we cannot lose data" тАФ and with asynchronous replication and automatic failover, that is false. A known 2012 GitHub incident: an out-of-date MySQL follower was promoted; it reused primary keys the old leader had already given out, and those keys were also used in a Redis store, so some private data was shown to the wrong users. The choice between sync and async is a choice between write latency and losing acknowledged writes, and you should make it on purpose.

ЁЯФз How (in this repo)

Leader(followers, sync=False) in dist/sim.py keeps a log and one list per follower. write(item) appends to the log and returns 'acknowledged' at once; with sync=True it also appends to every follower first. ship(upto) copies the first upto log entries to the followers (the async messenger). fail_over() promotes the follower with the longest copy and returns (its copy, the writes it lacks).

ЁЯзк Try it

python3 dist/demo.py replication
python3 - <<'EOF'
import sys; sys.path.insert(0, "dist"); from sim import Leader
for shipped in (0, 2, 3, 4):
    L = Leader(["nashik", "nagpur"])
    for g in ("g1", "g2", "g3", "g4"): L.write(g)
    L.ship(shipped)
    survivor, lost = L.fail_over()
    print(f"async, {shipped} of 4 shipped before the crash тЖТ new leader has {survivor}, lost {lost}")
EOF

тЬЕ Verify тАФ what you should see

replication prints:

тФАтФА asynchronous replication: 4 grades acknowledged, then Pune (leader) dies тЖТ new leader has ['g1', 'g2'], lost ['g3', 'g4']
тФАтФА synchronous replication: 4 grades acknowledged, then Pune (leader) dies тЖТ new leader has ['g1', 'g2', 'g3', 'g4'], lost []

Your snippet prints:

async, 0 of 4 shipped before the crash тЖТ new leader has [], lost ['g1', 'g2', 'g3', 'g4']
async, 2 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2'], lost ['g3', 'g4']
async, 3 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2', 'g3'], lost ['g4']
async, 4 of 4 shipped before the crash тЖТ new leader has ['g1', 'g2', 'g3', 'g4'], lost []

ЁЯПБ What you just proved

With async replication, what a failover loses is exactly the replication lag at the moment of the crash: 0 shipped тЖТ all 4 acknowledged grades lost; 4 shipped тЖТ none. The teacher heard "acknowledged" all four times. Synchronous replication lost nothing, because "acknowledged" waited for a copy.

тЪая╕П Common mistakes

ЁЯПн In production

PostgreSQL тАФ streaming replication is asynchronous by default. On a real account, make one of two standbys synchronous in postgresql.conf on the primary:

synchronous_standby_names = 'FIRST 1 (nashik, nagpur)'   # wait for the first one in the list that is connected
synchronous_commit = on                                  # a commit waits until that standby has flushed the WAL to disk

Or 'ANY 1 (nashik, nagpur)' тАФ wait for whichever answers first (a quorum commit). synchronous_commit also accepts remote_apply (wait until the standby has applied it, so reads there see it), remote_write, local and off, and it can be set per transaction тАФ for example SET LOCAL synchronous_commit = off; for a log line you can afford to lose. Watch the lag:

SELECT application_name, state, sync_state, write_lag, flush_lag, replay_lag
FROM pg_stat_replication;

MySQL тАФ semisynchronous replication is a plugin: the source waits until at least one replica has received the transaction (rpl_semi_sync_source_wait_for_replica_count).

Managed failover тАФ Amazon RDS Multi-AZ keeps a synchronous standby in another Availability Zone and fails over to it; RDS read replicas are asynchronous and do not fail over on their own. Tools such as Patroni run PostgreSQL failover using etcd or ZooKeeper (lessons 09тАУ10) to decide who the one leader is.

ЁЯПн Why this matters in production: for each database, write down: sync or async, how many synchronous copies, who decides a failover, and how the old leader is kept from taking writes. Then run a failover drill and count what was lost.

тПня╕П Next

A single leader is a bottleneck and a single point of decision. Another way: write to several copies, read from several, and let the numbers guarantee you see the newest. Quorums.

git checkout lesson-06-quorums
тЖР Previousfailure detectionNext тЖТquorums

This page is the lesson's README from the lesson-05-replication branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.