ЁЯПл The SchoolтА║ЁЯУР System DesignтА║ЁЯОм рдзрдбрд╛ 18 тАФ рд╕реЛрдбрд╡рд▓реЗрд▓реНрдпрд╛ рд░рдЪрдирд╛: file storage ┬╖ video ┬╖ RAG
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯОм рдзрдбрд╛ 18 тАФ рд╕реЛрдбрд╡рд▓реЗрд▓реНрдпрд╛ рд░рдЪрдирд╛: file storage ┬╖ video ┬╖ RAG

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 18 рдкреИрдХреА рдзрдбрд╛ 18 ┬╖ рдорд╛рдЧреЗ: lesson-17-designs-short-notify-chat


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рд╕рдВрдкреВрд░реНрдг course, рдЖрдгрд┐ рддреАрди рдЬрдб рд░рдЪрдирд╛, рдкреНрд░рддреНрдпреЗрдХ рддреНрдпрд╛рдЪ рдирдК рдкрд╛рдпрд▒реНрдпрд╛рдВрддреВрди: requirements тЖТ numbers тЖТ API тЖТ data тЖТ blocks тЖТ failure тЖТ security тЖТ observability тЖТ cost. File storage (chunks, content hashes, dedupe, metadata vs blobs, sync), video (upload тЖТ transcode ladder тЖТ CDN; egress рд╕рд░реНрд╡рд╛рдд рдореЛрдард╛) рдЖрдгрд┐ RAG (ingest offline, ask online; рдЖрдХрд╛рд░ рдЖрдгрд┐ token budgets).

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рджреАрдкрд┐рдХрд╛рдЪреНрдпрд╛ planning рдХрд╛рд░реНрдпрд╛рд▓рдпрд╛рдд рд╢реЗрд╡рдЯрдЪреНрдпрд╛ рддреАрди рдорд╛рдЧрдгреНрдпрд╛ рдпреЗрддрд╛рдд.

ЁЯУБ рд╕рд╛рдорд╛рдпрд┐рдХ files. рд╢рд┐рдХреНрд╖рд┐рдХрд╛ рд╢рд╛рд│реЗрдЪрд╛ рд╡рд╛рд░реНрд╖рд┐рдХ рдЕрд╣рд╡рд╛рд▓ рдПрдХрд╛ рд╕рд╛рдорд╛рдпрд┐рдХ folder рдордзреНрдпреЗ рдареЗрд╡рддрд╛рдд. рдПрдХ рд╢рдмреНрдж рдмрджрд▓рддреЛ. рдкреВрд░реНрдг file рдкреБрдиреНрд╣рд╛ рдкреНрд░рд╡рд╛рд╕ рдХрд░рд╛рдпрд▓рд╛ рд╣рд╡реА рдХрд╛? рдирд╛рд╣реА. рдХрд╛рд░реНрдпрд╛рд▓рдп рдкреНрд░рддреНрдпреЗрдХ file рдЪреЗ рдЫреЛрдЯреЗ рддреБрдХрдбреЗ рдХрд░рддреЗ рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рддреБрдХрдбреНрдпрд╛рд▓рд╛ рддреНрдпрд╛рдЪреНрдпрд╛ content рд╡рд░реВрди рдмрдирд▓реЗрд▓рд╛ рдард╕рд╛ рджреЗрддреЗ. рдлрдХреНрдд рдирд╡реНрдпрд╛ рдард╢рд╛рдЪреЗ рддреБрдХрдбреЗрдЪ рдкреНрд░рд╡рд╛рд╕ рдХрд░рддрд╛рдд. рджреЛрди рд╢рд┐рдХреНрд╖рд┐рдХрд╛рдВрдиреА рддреЛрдЪ form upload рдХреЗрд▓рд╛? рдХрд╛рд░реНрдпрд╛рд▓рдп рддреЛ рдПрдХрджрд╛рдЪ рдареЗрд╡рддреЗ.

ЁЯОм рд╡рд╛рд░реНрд╖рд┐рдХ рджрд┐рд╡рд╕рд╛рдЪреЗ videos. Upload рдХреЗрд▓реЗрд▓реЗ рдПрдХ рдорд┐рдирд┐рдЯ рдЪрд╛рд░ рдЖрдХрд╛рд░рд╛рдВрдд рдХреЙрдкреА рдХреЗрд▓реЗ рдЬрд╛рддреЗ, рд╕реНрдкрд╖реНрдЯрдкрд╛рд╕реВрди рд▓рд╣рд╛рдирдкрд░реНрдпрдВрдд, рдореНрд╣рдгрдЬреЗ рдкреНрд░рддреНрдпреЗрдХ phone рд▓рд╛ рддреНрдпрд╛рдЪреЗ network рд╡рд╛рд╣реВ рд╢рдХреЗрд▓ рдЕрд╕рд╛ рдЖрдХрд╛рд░ рдорд┐рд│рддреЛ. рддреЗ рд╕рд╛рдард╡рдгреЗ рд╕реНрд╡рд╕реНрдд рдЖрд╣реЗ. рддреЗ рд▓рд╛рдЦрднрд░ рдкрд╛рд▓рдХрд╛рдВрдирд╛ рдкрд╛рдард╡рдгреЗ рд╣реЗ рдореЛрдареЗ рдмрд┐рд▓ рдЖрд╣реЗ.

ЁЯдЦ рд╢рд╛рд│реЗрдЪреНрдпрд╛ documents рдирд╛ рд╡рд┐рдЪрд╛рд░рд╛. рдПрдХ рдкрд╛рд▓рдХ рд╡рд┐рдЪрд╛рд░рддреЗ: "рд╕рд╣рд▓реАрдЪреА рдлреА рдХрдзреА рднрд░рд╛рдпрдЪреА?" рдХрд╛рд░реНрдпрд╛рд▓рдпрд╛рдиреЗ рдЖрдзреАрдЪ 10,000 documents рдЪреЗ рдЫреЛрдЯреЗ рдкрд░рд┐рдЪреНрдЫреЗрдж рдХреЗрд▓реЗ рдЖрд╣реЗрдд рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рддреНрдпрд╛рдЪреНрдпрд╛ рдЕрд░реНрдерд╛рдиреБрд╕рд╛рд░ рд▓рд╛рд╡реВрди рдареЗрд╡рд▓рд╛ рдЖрд╣реЗ тАФ рд░рд╛рддреНрд░реА, рдХреЛрдгреА рд╡рд┐рдЪрд╛рд░рдгреНрдпрд╛рдЖрдзреА. рдкреНрд░рд╢реНрди рдЖрд▓рд╛ рдХреА рдХрд╛рд░реНрдпрд╛рд▓рдп рд╕рд░реНрд╡рд╛рдд рдЬрд╡рд│рдЪреЗ рдкрд╛рдЪ рдкрд░рд┐рдЪреНрдЫреЗрдж рд╢реЛрдзрддреЗ рдЖрдгрд┐ рдкреНрд░рд╢реНрдирд╛рд╕рд╣ рддреЗ AI рд▓рд╛ рджреЗрддреЗ. AI рдЙрддреНрддрд░ рджреЗрддреЗ рдЖрдгрд┐ рдЙрддреНрддрд░ рдХреБрдареЗ рд╕рд╛рдкрдбрд▓реЗ рддреЗ рджрд╛рдЦрд╡рддреЗ.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    subgraph fs["ЁЯУБ file storage"]
      f["file"] --> ck["тЬВя╕П chunks + SHA-256"] --> have{"server has it?"}
      have -->|"no"| s3[("ЁЯкг object storage")]
      md[("ЁЯЧДя╕П metadata DB<br/>file тЖТ chunk list")]
    end
    subgraph vd["ЁЯОм video"]
      up["тмЖя╕П upload"] --> tq["ЁЯУм queue"] --> tr["ЁЯОЮя╕П transcode<br/>1080p ┬╖ 720p ┬╖ 480p ┬╖ 240p"] --> cdn["ЁЯМН CDN<br/>egress = the bill"]
    end
    subgraph rg["ЁЯдЦ RAG"]
      ing["offline: chunk тЖТ embed тЖТ index"] --> vx[("ЁЯзн vector index")]
      q["online: question тЖТ embed"] --> vx --> top["top-5"] --> llm["ЁЯза prompt тЖТ answer + sources"]
    end

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/system-design/lesson-diagrams.html#l18

тЭУ рдХрд╛рдп

ЁЯУБ рд░рдЪрдирд╛ 4 тАФ file storage рдЖрдгрд┐ sync

рдкрд╛рдпрд░реА рдЖрд░рд╛рдЦрдбрд╛
requirements upload, download, share, devices рдордзреНрдпреЗ sync, versions; рдЕрдиреЗрдХ GB рдкрд░реНрдпрдВрддрдЪреНрдпрд╛ files; рд╡реНрдпрд╛рдкреНрддреАрдмрд╛рд╣реЗрд░: real time рдордзреНрдпреЗ рдПрдХрддреНрд░ editing
numbers file рдЪреНрдпрд╛ рдЖрдХрд╛рд░рд╛рдкреЗрдХреНрд╖рд╛ рдмрджрд▓рд╛рдЪрд╛ рдЖрдХрд╛рд░ рдЬрд╛рд╕реНрдд рдорд╣рддреНрддреНрд╡рд╛рдЪрд╛: рдПрдХрд╛ рд╢рдмреНрджрд╛рдЪреНрдпрд╛ рдмрджрд▓рд╛рд▓рд╛ рдПрдХрд╛ chunk рдЪреА рдХрд┐рдВрдордд рд▓рд╛рдЧрд╛рд╡реА, рдкреВрд░реНрдг file рдЪреА рдирд╛рд╣реА
API POST /files/{id}/versions {chunk_hashes} тЖТ {missing} ┬╖ PUT /chunks/{hash} ┬╖ GET /files/{id} тЖТ chunk list
data database рдордзреНрдпреЗ metadata: files, versions(file_id, version, chunk_hashes[]), shares; object storage рдордзреНрдпреЗ blobs (chunks), hash рд╣реА key
blocks client рдмрд╛рдЬреВрдЪрд╛ chunker; pre-signed URLs рдиреЗ рдереЗрдЯ object storage рд╡рд░ upload; рдирд╡реА version рдЖрд▓реА рдХреА рдЗрддрд░ devices рдирд╛ notification
failure uploads рдкреНрд░рддреНрдпреЗрдХ chunk рдкрд╛рд╕реВрди рдкреБрдиреНрд╣рд╛ рд╕реБрд░реВ рд╣реЛрддрд╛рдд; рдкреНрд░рддреНрдпреЗрдХ chunk рд╕рд╛рдард╡рд▓реНрдпрд╛рд╡рд░рдЪ version рджрд┐рд╕рддреЗ (metadata рд╢реЗрд╡рдЯреА commit рдХрд░рд╛); conflicts тЖТ рджреЛрдиреНрд╣реА рдареЗрд╡рд╛ ("conflicted copy")
security рдкреНрд░рддреНрдпреЗрдХ file рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ share link рд╕рд╛рдареА authorization; at rest encryption; рдлрдХреНрдд chunk hash рдиреЗ access рдорд┐рд│реВ рдирдпреЗ
observability upload success rate, dedupe рдиреЗ рд╡рд╛рдЪрд▓реЗрд▓реЗ bytes, devices рдордзрд▓рд╛ sync delay
cost storage рдкреНрд░рддреАрдВрдиреА рдирд╛рд╣реА, unique chunks рдиреЗ рд╡рд╛рдврддреЗ тАФ dedupe рдореНрд╣рдгрдЬреЗ рдкреИрд╕рд╛

ЁЯОм рд░рдЪрдирд╛ 5 тАФ video

рдкрд╛рдпрд░реА рдЖрд░рд╛рдЦрдбрд╛
requirements video upload рдХрд░рд╛; рдХреЛрдгрддреНрдпрд╛рд╣реА phone рдЖрдгрд┐ network рд╡рд░ рдкрд╛рд╣рд╛; рд╕рд╛рдзрд╛рд░рдг 2 рд╕реЗрдХрдВрджрд╛рдВрдд play рд╕реБрд░реВ
numbers рдПрдХ рдорд┐рдирд┐рдЯ тЖТ 4 renditions тЖТ 0.07 GB рд╕рд╛рдард╡рд▓реЗ; 1M views ├Ч 10 min, 2.5 Mb/s рдиреЗ тЖТ 187.5 TB egress
API POST /videos тЖТ рдПрдХ pre-signed upload URL ┬╖ GET /videos/{id} тЖТ рдПрдХ playlist URL (HLS / DASH)
data videos(id, owner, status, duration, renditions); files object storage рдордзреНрдпреЗ
blocks upload тЖТ object storage тЖТ event тЖТ queue тЖТ transcode workers тЖТ segments + playlists тЖТ CDN
failure transcoding рд╣рд╛ queue рдордзрд▓рд╛, retry рд╣реЛрдгрд╛рд░рд╛, idempotent job рдЖрд╣реЗ (рдзрдбрд╛ 07); рдПрдХ rendition рдЕрдкрдпрд╢реА рдЭрд╛рд▓реА рддрд░ рдмрд╛рдХреАрдЪреНрдпрд╛ рдЕрдбрдд рдирд╛рд╣реАрдд
security рдлрдХреНрдд рд╢рд╛рд│реЗрдЪреА рдХреБрдЯреБрдВрдмреЗрдЪ рдкрд╛рд╣реВ рд╢рдХрддрд╛рдд: CDN рд╡рд░ signed URLs рдХрд┐рдВрд╡рд╛ signed cookies; upload рд╡рд░ malware рдЖрдгрд┐ content рддрдкрд╛рд╕рдгреА
observability рдкрд╣рд┐рд▓реНрдпрд╛ frame рдкрд░реНрдпрдВрддрдЪрд╛ рд╡реЗрд│, rebuffering ratio, transcode queue depth, CDN cache hit ratio
cost egress рд╕рд░реНрд╡рд╛рдд рдореЛрдард╛: рдЙрджрд╛рд╣рд░рдгрд╛рджрд╛рдЦрд▓ рдХрд┐рдорддреАрд▓рд╛ 187.5 TB рдкрд╛рдард╡рд╛рдпрд▓рд╛ рд╕рд╛рдзрд╛рд░рдг $16,875; 1,000 рддрд╛рд╕ рд╕рд╛рдард╡рд╛рдпрд▓рд╛ рдорд╣рд┐рдиреНрдпрд╛рд▓рд╛ рд╕рд╛рдзрд╛рд░рдг $91

ЁЯдЦ рд░рдЪрдирд╛ 6 тАФ RAG (retrieval-augmented generation)

рдкрд╛рдпрд░реА рдЖрд░рд╛рдЦрдбрд╛
requirements рдкрд╛рд▓рдХ рд╢рд╛рд│реЗрдЪреНрдпрд╛ documents рдмрджреНрджрд▓ рдкреНрд░рд╢реНрди рд╡рд┐рдЪрд╛рд░рддрд╛рдд; рдЙрддреНрддрд░реЗ рдЖрдкрд▓реЗ sources рд╕рд╛рдВрдЧрддрд╛рдд; рдЬреНрдпрд╛рдВрдирд╛ рдкрд╛рд╣рд╛рдпрдЪреА рдкрд░рд╡рд╛рдирдЧреА рдирд╛рд╣реА рддреЗ documents рдХреЛрдгрд╛рд▓рд╛рдЪ рджрд┐рд╕рдд рдирд╛рд╣реАрдд
numbers 10,000 documents ├Ч 20 pages ├Ч 500 words, рдкреНрд░рддреНрдпреЗрдХ chunk рдордзреНрдпреЗ 250 words тЖТ 400,000 chunks; 1,024 dimensions ├Ч 4 bytes тЖТ 1.64 GB index; рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рд▓рд╛ ~1,675 prompt tokens
API POST /ask {question} тЖТ {answer, sources[]}
data chunks(id, doc_id, text, embedding, allowed_classes); documents object storage рдордзреНрдпреЗ
blocks ingest (offline, queue рдордзреВрди): parse тЖТ chunk тЖТ embed тЖТ index. Ask (online): рдкреНрд░рд╢реНрди embed рдХрд░рд╛ тЖТ permission filter рд╕рд╣ top-k search тЖТ prompt тЖТ sources рд╕рд╣ рдЙрддреНрддрд░
failure model рд╣рд│реВ рдХрд┐рдВрд╡рд╛ рдмрдВрдж тЖТ AI рдЙрддреНрддрд░рд╛рд╢рд┐рд╡рд╛рдп top passages рджрд╛рдЦрд╡рд╛ (degrade); рдкреНрд░рддреНрдпреЗрдХ request рд╕рд╛рдареА timeouts рдЖрдгрд┐ token budget
security permission рдиреБрд╕рд╛рд░ filter search рдордзреНрдпреЗрдЪ рдХрд░рд╛, рдирдВрддрд░ рдирд╛рд╣реА; document рдЪрд╛ рдордЬрдХреВрд░ data рдореНрд╣рдгреВрди рд╣рд╛рддрд╛рд│рд╛, рдХрдзреАрд╣реА instructions рдореНрд╣рдгреВрди рдирд╛рд╣реА (prompt injection); рддреБрдордЪреНрдпрд╛ рдирд┐рдпрдВрддреНрд░рдгрд╛рдмрд╛рд╣реЗрд░ рдЬрд╛рдгрд╛рд▒реНрдпрд╛ prompts рдордзреНрдпреЗ рд╡реИрдпрдХреНрддрд┐рдХ data рдирдХреЛ
observability retrieval quality (рдпреЛрдЧреНрдп passage рдкрд░рдд рдЖрд▓рд╛ рдХрд╛?), source рд╕рд╣ рдЙрддреНрддрд░реЗ, рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рдЪреЗ tokens, рдкреВрд░реНрдг ask рдЪрд╛ p99
cost embeddings рд╕рд╛рдареА ingest рд╡реЗрд│реА рдПрдХрджрд╛рдЪ рдкреИрд╕реЗ; tokens рд╕рд╛рдареА рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рд▓рд╛ рдкреИрд╕реЗ тАФ prompt рдЪрд╛ рдЖрдХрд╛рд░ рд╣рд╛ рдЪрд╛рд▓реВ рдЦрд░реНрдЪ

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг рдЬрдб data рдореБрд│реЗ рд░рдЪрдиреЗрдЪреА рдХреЛрдгрддреА рдУрд│ рдорд╣рддреНрддреНрд╡рд╛рдЪреА рддреЗ рдмрджрд▓рддреЗ. Files рд╕рд╛рдареА рдмрджрд▓рд╛рдЪрд╛ рдЖрдХрд╛рд░ bandwidth рдард░рд╡рддреЛ, рдореНрд╣рдгреВрди content hashes рд╣рд╛ рдЧрд╛рднрд╛. Video рд╕рд╛рдареА egress рдмрд┐рд▓ рдард░рд╡рддреЛ, рдореНрд╣рдгреВрди CDN рдЖрдгрд┐ ladder рд╣рд╛ рдЧрд╛рднрд╛. RAG рд╕рд╛рдареА offline/online рд╡рд┐рднрд╛рдЧрдгреА latency рдард░рд╡рддреЗ, рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рдЪреЗ tokens рдЪрд╛рд▓реВ рдЦрд░реНрдЪ рдард░рд╡рддрд╛рдд. рддреНрдпрд╛рдЪ рдирдК рдкрд╛рдпрд▒реНрдпрд╛ рдпрд╛рддрд▓реЗ рдкреНрд░рддреНрдпреЗрдХ рд╢реЛрдзрддрд╛рдд тАФ рдкрджреНрдзрддреАрдЪрд╛ рд╣рд╛рдЪ рддрд░ рдореБрджреНрджрд╛ рдЖрд╣реЗ.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

design/designs.py рдордзреНрдпреЗ:

design/demo.py рдордзрд▓реЗ designs2() рддрд┐рдиреНрд╣реА рдЪрд╛рд▓рд╡рддреЗ. snippet рдзрдбрд╛ 16 рдордзрд▓реНрдпрд╛ monthly_cost() рдиреЗ video рдЪреА рдХрд┐рдВрдорддрд╣реА рдХрд╛рдврддреЛ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 design/demo.py designs2
python3 - <<'EOF'
import sys; sys.path.insert(0, "design"); from designs import chunks, sync_upload, video_storage_gb, cdn_egress_tb, rag, monthly_cost
old = "the school annual report for 2026 version one"
for new, what in ((old, "nothing changed"), (old.replace("one", "two"), "one word replaced"), ("A " + old, "two letters added at the start")):
    up, total = sync_upload(old, new); print(f"{what:<32} тЖТ upload {up:>2} of {total} chunks")
print("same content, same chunk id:", chunks("dipika") == chunks("dipika"), chunks("dipika"))
for minutes in (1, 10, 60):
    print(f"{minutes:>2} min uploaded тЖТ {video_storage_gb(minutes):>5} GB stored (all renditions)")
tb = cdn_egress_tb(1_000_000, 10)
print(f"1M views ├Ч 10 min тЖТ {tb} TB egress ┬╖ at the example $0.09/GB тЖТ", monthly_cost(0, 0, tb * 1000, 0)["egress"], "$ ┬╖ storing 1,000 hours тЖТ", monthly_cost(0, video_storage_gb(60) * 1000, 0, 0)["storage"], "$ a month")
for docs, k, words in ((10_000, 5, 250), (10_000, 10, 250), (100_000, 5, 250), (10_000, 5, 500)):
    print(f"{docs:>7,} docs, top-{k:<2}, {words} words per chunk тЖТ {rag(docs, 20, chunk_words=words, top_k=k)}")
EOF
python3 design/test_design.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

designs2 рд╣реЗ рдЫрд╛рдкрддреЗ:

тФАтФА file storage: a file split into chunks by content hash; one word changed тЖТ upload 2 of 12 chunks
   metadata (who, which chunks, which version) in a database ┬╖ chunks in object storage ┬╖ dedupe for free
тФАтФА video: one uploaded minute becomes 4 renditions (1080p, 720p, 480p, 240p) тЖТ 0.07 GB per minute stored
   1M views ├Ч 10 minutes at 2.5 Mb/s тЖТ 187.5 TB from the CDN тАФ egress, not storage, is the bill
тФАтФА RAG over 10,000 school documents ├Ч 20 pages тЖТ 400,000 chunks ┬╖ vector index ~1.64 GB ┬╖ ~1675 prompt tokens per question
   ingest (chunk тЖТ embed тЖТ index) runs offline; ask (embed question тЖТ top-5 тЖТ prompt тЖТ answer with sources) runs online
тФАтФА the method, every time: requirements тЖТ numbers тЖТ API тЖТ data тЖТ blocks тЖТ failure тЖТ security тЖТ observability тЖТ cost тЖТ ADR

рддреБрдордЪрд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

nothing changed                  тЖТ upload  0 of 12 chunks
one word replaced                тЖТ upload  2 of 12 chunks
two letters added at the start   тЖТ upload 12 of 12 chunks
same content, same chunk id: True ['0f12ca8a', '3ebb1c99']
 1 min uploaded тЖТ  0.07 GB stored (all renditions)
10 min uploaded тЖТ  0.66 GB stored (all renditions)
60 min uploaded тЖТ  3.96 GB stored (all renditions)
1M views ├Ч 10 min тЖТ 187.5 TB egress ┬╖ at the example $0.09/GB тЖТ 16875.0 $ ┬╖ storing 1,000 hours тЖТ 91.08 $ a month
 10,000 docs, top-5 , 250 words per chunk тЖТ {'chunks': 400000, 'index_gb': 1.64, 'prompt_tokens': 1675}
 10,000 docs, top-10, 250 words per chunk тЖТ {'chunks': 400000, 'index_gb': 1.64, 'prompt_tokens': 3300}
100,000 docs, top-5 , 250 words per chunk тЖТ {'chunks': 4000000, 'index_gb': 16.38, 'prompt_tokens': 1675}
 10,000 docs, top-5 , 500 words per chunk тЖТ {'chunks': 200000, 'index_gb': 0.82, 'prompt_tokens': 3300}

Tests тЬЕ L18 changing one word re-uploads only the changed chunks рдЖрдгрд┐ 18/18 passed рдиреЗ рд╕рдВрдкрддрд╛рдд.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

Content hashes рдореБрд│реЗ sync рд╕реНрд╡рд╕реНрдд рд╣реЛрддреЗ: рдмрджрд▓ рдирд╛рд╣реА тЖТ 0 chunks, рдПрдХ рд╢рдмреНрдж тЖТ 12 рдкреИрдХреА 2. рдкрдг fixed-size chunks рдЪреА рдПрдХ рдХрдордЬреЛрд░реА рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдкрд╛рд╣рд┐рд▓реА: рд╕реБрд░реБрд╡рд╛рддреАрдЪреНрдпрд╛ рджреЛрди рдЕрдХреНрд╖рд░рд╛рдВрдиреА рдкреНрд░рддреНрдпреЗрдХ chunk рд╕рд░рдХрд╡рд▓рд╛ (12 рдкреИрдХреА 12) тАФ рдореНрд╣рдгреВрдирдЪ рдЦрд░реА tools content-defined chunks рд╡рд╛рдкрд░рддрд╛рдд. Video рд╕рд╛рдареА, рдПрдХрд╛ рдорд╣рд┐рдиреНрдпрд╛рдЪреЗ views рдкрд╛рдард╡рдгреЗ ($16,875) рд╣реЗ 1,000 рддрд╛рд╕ рд╕рд╛рдард╡рдгреНрдпрд╛рдкреЗрдХреНрд╖рд╛ ($91.08) рд╕рд╛рдзрд╛рд░рдг 185 рдкрдЯ рдорд╣рд╛рдЧ рдЖрд╣реЗ: рдЖрдзреА CDN рдЖрдгрд┐ ladder рдЖрдЦрд╛. RAG рд╕рд╛рдареА, рджрд╣рд╛ рдкрдЯ documents рдореНрд╣рдгрдЬреЗ рджрд╣рд╛ рдкрдЯ index (16.38 GB) рдкрдг prompt рддреЛрдЪ (1,675 tokens) тАФ рддрд░ top-k рдХрд┐рдВрд╡рд╛ chunk size рджреБрдкреНрдкрдЯ рдХреЗрд▓реНрдпрд╛рд╕ рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рд▓рд╛ рдореЛрдЬрд╛рд╡реЗ рд▓рд╛рдЧрдгрд╛рд░реЗ tokens рджреБрдкреНрдкрдЯ рд╣реЛрддрд╛рдд.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ рдереЗрдЯ S3 рд╡рд░ multipart upload (рдкреНрд░рддреНрдпреЗрдХ part рдПрдХрдЯреНрдпрд╛рдиреЗ retry рд╣реЛрддреЛ; parts рд╕рдорд╛рдВрддрд░ рдЬрд╛рдК рд╢рдХрддрд╛рдд):

aws s3api create-multipart-upload --bucket school-files --key chunks/3ebb1c99
aws s3api upload-part --bucket school-files --key chunks/3ebb1c99 --part-number 1 \
    --body part-1.bin --upload-id "EXAMPLE-UPLOAD-ID"
aws s3api complete-multipart-upload --bucket school-files --key chunks/3ebb1c99 \
    --upload-id "EXAMPLE-UPLOAD-ID" --multipart-upload file://parts.json

рдПрдХ pre-signed URL, рдореНрд╣рдгрдЬреЗ phone рддреБрдордЪреНрдпрд╛ API рдордзреВрди рди рдЬрд╛рддрд╛ object storage рд╡рд░ upload рдХрд░рддреЛ:

aws s3 presign s3://school-videos/uploads/annual-day.mp4 --expires-in 900

Video: AWS Elemental MediaConvert job upload рдЪреЗ HLS ladder рдордзреНрдпреЗ рд░реВрдкрд╛рдВрддрд░ рдХрд░рддреЛ (job settings file рдордзреНрдпреЗ рдЪрд╛рд░ renditions рдЕрд╕рддрд╛рдд), рдЖрдгрд┐ CloudFront segments рдкреБрд░рд╡рддреЛ:

aws mediaconvert create-job --endpoint-url https://abcd1234.mediaconvert.ap-south-1.amazonaws.com \
    --role arn:aws:iam::111122223333:role/mediaconvert-school \
    --settings file://hls-ladder-1080-720-480-240.json

pgvector рд╕рд╣ PostgreSQL рд╡рд░ RAG: HNSW index, рдЖрдгрд┐ permission filter query рдЪреНрдпрд╛ рдЖрдд:

CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (id bigserial PRIMARY KEY, doc_id bigint, body text,
                     allowed_classes text[], embedding vector(1024));
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

SELECT doc_id, body
FROM chunks
WHERE allowed_classes && ARRAY['3A']               -- only what this parent may see
ORDER BY embedding <=> $1                           -- $1 = the question's embedding
LIMIT 5;

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдкреНрд░рддреНрдпреЗрдХ рдЬрдб рд░рдЪрдиреЗрд╕рд╛рдареА рдмрд┐рд▓рд╛рд╢реЗрдЬрд╛рд░реА dashboard рд╡рд░ рдПрдХ рдЖрдХрдбрд╛ рдареЗрд╡рд╛: dedupe рдиреЗ рд╡рд╛рдЪрд▓реЗрд▓реЗ bytes, рджрд░рд░реЛрдЬрдЪрд╛ CDN egress, рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рд╢реНрдирд╛рдЪреЗ tokens. рддреЗ рдХреЛрдгрддреНрдпрд╛рд╣реА server рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд рдЦрд░реНрдЪ рд╣рд▓рд╡рддрд╛рдд.

ЁЯОУ рдЖрддрд╛ planning рдХрд╛рд░реНрдпрд╛рд▓рдп рддреБрдордЪреЗ рдЖрд╣реЗ

рдХрд╛рдврдгреНрдпрд╛рдЖрдзреА рд╡рд┐рдЪрд╛рд░рд╛ тЖТ рдореЛрдЬрд╛ тЖТ рдЖрдзреА рдХрд░рд╛рд░ тЖТ рдкреНрд░рд╢реНрдирд╛рдВрд╡рд░реВрди data тЖТ рдкреНрд░рддреА тЖТ рдкрд╕рд░рд╡рдгреЗ тЖТ рдирдВрддрд░ тЖТ рдЕрдиреЗрдХ рдРрдХрдгрд╛рд░реЗ рдЕрд╕рд▓реЗрд▓реА рддрдереНрдпреЗ тЖТ рдПрдХ deploy рдХреА рдЕрдиреЗрдХ тЖТ рдкреНрд░рддреАрдВрдордзрд▓реЗ рдПрдХрдордд тЖТ рдкреНрд░рддреНрдпреЗрдХ рдкрд╛рдпрд░реАрд▓рд╛ рдПрдХ undo тЖТ рдмрд░рдгреАрддрд▓реЗ tokens тЖТ рдкреНрд░рддреНрдпреЗрдХ рдмрд╛рдгрд╛рд▓рд╛ рдПрдХ рдпреЛрдЬрдирд╛ тЖТ рдкреНрд░рддреНрдпреЗрдХ box рд╕рд╛рдареА рд╕рд╣рд╛ рдкреНрд░рд╢реНрди тЖТ рдирд┐рд░реЛрдЧреА рдореНрд╣рдгрдЬреЗ рдХрд╛рдп тЖТ рдПрдХ рдмрд┐рд▓ рдЖрдгрд┐ рдПрдХ рдХрд╛рд░рдг тЖТ рд╕рд╣рд╛ рд░рдЪрдирд╛, рд╕реБрд░реБрд╡рд╛рддреАрдкрд╛рд╕реВрди рд╢реЗрд╡рдЯрдкрд░реНрдпрдВрдд. рддреБрдореНрд╣реА рдлрдХреНрдд system design рд╢рд┐рдХрд▓рд╛ рдирд╛рд╣реАрдд тАФ рдХреЛрдгреА рдмрд╛рдВрдзрдгреНрдпрд╛рдЖрдзреАрдЪ рддреБрдореНрд╣реА system рдЖрдЦреВ рд╢рдХрддрд╛, рдЖрдгрд┐ рднрд┐рдВрддреАрд╡рд░рдЪрд╛ рдкреНрд░рддреНрдпреЗрдХ box рд╕рдордЬрд╛рд╡реВ рд╢рдХрддрд╛. ЁЯУРЁЯПлЁЯОУ

тПня╕П рдкреБрдвреЗ

рдЗрддрд░ рд╢рд╛рд│рд╛, рдЬрд┐рдереЗ рддреБрдордЪреНрдпрд╛ рдЖрд░рд╛рдЦрдбреНрдпрд╛рддрд▓реЗ boxes рдмрд╛рдВрдзрд▓реЗ рдЬрд╛рддрд╛рдд:

git checkout main
python3 design/demo.py     # one last run, for fun

ЁЯОм Lesson 18 тАФ Worked designs: file storage ┬╖ video ┬╖ RAG

ЁЯУН You are here: Lesson 18 of 18 ┬╖ Previous: lesson-17-designs-short-notify-chat


ЁЯУж What's in this branch

The whole course, plus three heavy designs, each walked through the same nine steps: requirements тЖТ numbers тЖТ API тЖТ data тЖТ blocks тЖТ failure тЖТ security тЖТ observability тЖТ cost. File storage (chunks, content hashes, dedupe, metadata vs blobs, sync), video (upload тЖТ transcode ladder тЖТ CDN; egress dominates) and RAG (ingest offline, ask online; sizes and token budgets).

ЁЯзТ Explain like I'm 5

Three last requests come to Dipika's planning office.

ЁЯУБ Shared files. Teachers keep the school's annual report in a shared folder. One word changes. Must the whole file travel again? No. The office cuts every file into small pieces and gives each piece a fingerprint made from its content. Only pieces with a new fingerprint travel. Two teachers upload the same form? The office keeps it once.

ЁЯОм Videos of the annual day. One uploaded minute is copied into four sizes, from sharp to small, so every phone gets a size its network can carry. Storing them is cheap. Sending them to a million parents is the big bill.

ЁЯдЦ Ask the school documents. A parent asks: "When is the fee due for the trip?" The office has already cut 10,000 documents into small passages and filed each one by its meaning тАФ at night, before anyone asks. When the question comes, the office finds the five closest passages and gives them to the AI with the question. The AI answers and shows where it found the answer.

ЁЯЧ║я╕П Diagram

flowchart LR
    subgraph fs["ЁЯУБ file storage"]
      f["file"] --> ck["тЬВя╕П chunks + SHA-256"] --> have{"server has it?"}
      have -->|"no"| s3[("ЁЯкг object storage")]
      md[("ЁЯЧДя╕П metadata DB<br/>file тЖТ chunk list")]
    end
    subgraph vd["ЁЯОм video"]
      up["тмЖя╕П upload"] --> tq["ЁЯУм queue"] --> tr["ЁЯОЮя╕П transcode<br/>1080p ┬╖ 720p ┬╖ 480p ┬╖ 240p"] --> cdn["ЁЯМН CDN<br/>egress = the bill"]
    end
    subgraph rg["ЁЯдЦ RAG"]
      ing["offline: chunk тЖТ embed тЖТ index"] --> vx[("ЁЯзн vector index")]
      q["online: question тЖТ embed"] --> vx --> top["top-5"] --> llm["ЁЯза prompt тЖТ answer + sources"]
    end

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/system-design/lesson-diagrams.html#l18

тЭУ What

ЁЯУБ Design 4 тАФ file storage and sync

step the plan
requirements upload, download, share, sync across devices, versions; files up to many GB; out of scope: editing together in real time
numbers the size of change matters more than file size: a one-word change should cost one chunk, not the file
API POST /files/{id}/versions {chunk_hashes} тЖТ {missing} ┬╖ PUT /chunks/{hash} ┬╖ GET /files/{id} тЖТ chunk list
data metadata in a database: files, versions(file_id, version, chunk_hashes[]), shares; blobs (the chunks) in object storage, keyed by hash
blocks client-side chunker; upload directly to object storage with pre-signed URLs; a notification to other devices when a new version exists
failure uploads resume per chunk; a version is visible only when every chunk is stored (commit the metadata last); conflicts тЖТ keep both ("conflicted copy")
security authorization per file and per share link; encryption at rest; a chunk hash alone must not grant access
observability upload success rate, bytes saved by dedupe, sync delay between devices
cost storage grows with unique chunks, not with copies тАФ dedupe is money

ЁЯОм Design 5 тАФ video

step the plan
requirements upload a video; watch on any phone and network; start playing in about 2 seconds
numbers one minute тЖТ 4 renditions тЖТ 0.07 GB stored; 1M views ├Ч 10 min at 2.5 Mb/s тЖТ 187.5 TB egress
API POST /videos тЖТ a pre-signed upload URL ┬╖ GET /videos/{id} тЖТ a playlist URL (HLS / DASH)
data videos(id, owner, status, duration, renditions); the files in object storage
blocks upload тЖТ object storage тЖТ event тЖТ queue тЖТ transcode workers тЖТ segments + playlists тЖТ CDN
failure transcoding is a queued, retried, idempotent job (lesson 07); a failed rendition does not block the others
security only the school's families may watch: signed URLs or signed cookies on the CDN; malware and content checks on upload
observability time to first frame, rebuffering ratio, transcode queue depth, CDN cache hit ratio
cost egress dominates: sending 187.5 TB costs about $16,875 at the example price; storing 1,000 hours costs about $91 a month

ЁЯдЦ Design 6 тАФ RAG (retrieval-augmented generation)

step the plan
requirements parents ask questions about school documents; answers cite their sources; nobody sees documents they may not see
numbers 10,000 documents ├Ч 20 pages ├Ч 500 words, 250 words per chunk тЖТ 400,000 chunks; 1,024 dimensions ├Ч 4 bytes тЖТ 1.64 GB index; ~1,675 prompt tokens per question
API POST /ask {question} тЖТ {answer, sources[]}
data chunks(id, doc_id, text, embedding, allowed_classes); the documents in object storage
blocks ingest (offline, queued): parse тЖТ chunk тЖТ embed тЖТ index. Ask (online): embed the question тЖТ top-k search with a permission filter тЖТ prompt тЖТ answer with sources
failure the model is slow or down тЖТ show the top passages without an AI answer (degrade); timeouts and a token budget per request
security filter by permission in the search, not after; treat document text as data, never as instructions (prompt injection); no personal data in prompts that leave your control
observability retrieval quality (did the right passage come back?), answers with a source, tokens per question, p99 of the whole ask
cost embeddings are paid once at ingest; tokens are paid on every question тАФ the prompt size is the running cost

ЁЯдФ Why

Because heavy data changes which line of the design matters. For files, the size of a change decides the bandwidth, so content hashes are the core. For video, egress decides the bill, so the CDN and the ladder are the core. For RAG, the offline/online split decides the latency, and tokens per question decide the running cost. The same nine steps find each of these тАФ which is the point of a method.

ЁЯФз How (in this repo)

In design/designs.py:

designs2() in design/demo.py runs all three. The snippet also prices video with monthly_cost() from lesson 16.

ЁЯзк Try it

python3 design/demo.py designs2
python3 - <<'EOF'
import sys; sys.path.insert(0, "design"); from designs import chunks, sync_upload, video_storage_gb, cdn_egress_tb, rag, monthly_cost
old = "the school annual report for 2026 version one"
for new, what in ((old, "nothing changed"), (old.replace("one", "two"), "one word replaced"), ("A " + old, "two letters added at the start")):
    up, total = sync_upload(old, new); print(f"{what:<32} тЖТ upload {up:>2} of {total} chunks")
print("same content, same chunk id:", chunks("dipika") == chunks("dipika"), chunks("dipika"))
for minutes in (1, 10, 60):
    print(f"{minutes:>2} min uploaded тЖТ {video_storage_gb(minutes):>5} GB stored (all renditions)")
tb = cdn_egress_tb(1_000_000, 10)
print(f"1M views ├Ч 10 min тЖТ {tb} TB egress ┬╖ at the example $0.09/GB тЖТ", monthly_cost(0, 0, tb * 1000, 0)["egress"], "$ ┬╖ storing 1,000 hours тЖТ", monthly_cost(0, video_storage_gb(60) * 1000, 0, 0)["storage"], "$ a month")
for docs, k, words in ((10_000, 5, 250), (10_000, 10, 250), (100_000, 5, 250), (10_000, 5, 500)):
    print(f"{docs:>7,} docs, top-{k:<2}, {words} words per chunk тЖТ {rag(docs, 20, chunk_words=words, top_k=k)}")
EOF
python3 design/test_design.py

тЬЕ Verify тАФ what you should see

designs2 prints:

тФАтФА file storage: a file split into chunks by content hash; one word changed тЖТ upload 2 of 12 chunks
   metadata (who, which chunks, which version) in a database ┬╖ chunks in object storage ┬╖ dedupe for free
тФАтФА video: one uploaded minute becomes 4 renditions (1080p, 720p, 480p, 240p) тЖТ 0.07 GB per minute stored
   1M views ├Ч 10 minutes at 2.5 Mb/s тЖТ 187.5 TB from the CDN тАФ egress, not storage, is the bill
тФАтФА RAG over 10,000 school documents ├Ч 20 pages тЖТ 400,000 chunks ┬╖ vector index ~1.64 GB ┬╖ ~1675 prompt tokens per question
   ingest (chunk тЖТ embed тЖТ index) runs offline; ask (embed question тЖТ top-5 тЖТ prompt тЖТ answer with sources) runs online
тФАтФА the method, every time: requirements тЖТ numbers тЖТ API тЖТ data тЖТ blocks тЖТ failure тЖТ security тЖТ observability тЖТ cost тЖТ ADR

Your snippet prints:

nothing changed                  тЖТ upload  0 of 12 chunks
one word replaced                тЖТ upload  2 of 12 chunks
two letters added at the start   тЖТ upload 12 of 12 chunks
same content, same chunk id: True ['0f12ca8a', '3ebb1c99']
 1 min uploaded тЖТ  0.07 GB stored (all renditions)
10 min uploaded тЖТ  0.66 GB stored (all renditions)
60 min uploaded тЖТ  3.96 GB stored (all renditions)
1M views ├Ч 10 min тЖТ 187.5 TB egress ┬╖ at the example $0.09/GB тЖТ 16875.0 $ ┬╖ storing 1,000 hours тЖТ 91.08 $ a month
 10,000 docs, top-5 , 250 words per chunk тЖТ {'chunks': 400000, 'index_gb': 1.64, 'prompt_tokens': 1675}
 10,000 docs, top-10, 250 words per chunk тЖТ {'chunks': 400000, 'index_gb': 1.64, 'prompt_tokens': 3300}
100,000 docs, top-5 , 250 words per chunk тЖТ {'chunks': 4000000, 'index_gb': 16.38, 'prompt_tokens': 1675}
 10,000 docs, top-5 , 500 words per chunk тЖТ {'chunks': 200000, 'index_gb': 0.82, 'prompt_tokens': 3300}

The tests end with тЬЕ L18 changing one word re-uploads only the changed chunks and 18/18 passed.

ЁЯПБ What you just proved

Content hashes make sync cheap: no change тЖТ 0 chunks, one word тЖТ 2 of 12. But fixed-size chunks have a weakness you just saw: two letters at the start shifted every chunk (12 of 12) тАФ that is why real tools use content-defined chunks. For video, sending one month of views ($16,875) costs about 185 times more than storing 1,000 hours ($91.08): design the CDN and the ladder first. For RAG, ten times the documents gives ten times the index (16.38 GB) but the same prompt (1,675 tokens) тАФ while doubling top-k or the chunk size doubles the tokens paid on every question.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ a multipart upload straight to S3 (each part is retried alone; parts can go in parallel):

aws s3api create-multipart-upload --bucket school-files --key chunks/3ebb1c99
aws s3api upload-part --bucket school-files --key chunks/3ebb1c99 --part-number 1 \
    --body part-1.bin --upload-id "EXAMPLE-UPLOAD-ID"
aws s3api complete-multipart-upload --bucket school-files --key chunks/3ebb1c99 \
    --upload-id "EXAMPLE-UPLOAD-ID" --multipart-upload file://parts.json

A pre-signed URL, so the phone uploads to object storage without going through your API:

aws s3 presign s3://school-videos/uploads/annual-day.mp4 --expires-in 900

Video: an AWS Elemental MediaConvert job turns the upload into an HLS ladder (the job settings file holds the four renditions), and CloudFront serves the segments:

aws mediaconvert create-job --endpoint-url https://abcd1234.mediaconvert.ap-south-1.amazonaws.com \
    --role arn:aws:iam::111122223333:role/mediaconvert-school \
    --settings file://hls-ladder-1080-720-480-240.json

RAG on PostgreSQL with pgvector: an HNSW index, and the permission filter inside the query:

CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (id bigserial PRIMARY KEY, doc_id bigint, body text,
                     allowed_classes text[], embedding vector(1024));
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

SELECT doc_id, body
FROM chunks
WHERE allowed_classes && ARRAY['3A']               -- only what this parent may see
ORDER BY embedding <=> $1                           -- $1 = the question's embedding
LIMIT 5;

ЁЯПн Why this matters in production: for each heavy design, keep one number on a dashboard next to the bill: bytes saved by dedupe, CDN egress per day, tokens per question. They move the cost more than any server.

ЁЯОУ The planning office is yours

Ask before you draw тЖТ count тЖТ the contract first тЖТ data from its questions тЖТ copies тЖТ spreading тЖТ later тЖТ facts with many listeners тЖТ one deploy or many тЖТ agreement between copies тЖТ an undo for every step тЖТ tokens in a jar тЖТ a plan for every arrow тЖТ six questions for every box тЖТ what healthy means тЖТ a bill and a reason тЖТ six designs, end to end. You didn't just learn system design тАФ you can plan a system before anyone builds it, and explain every box on the wall. ЁЯУРЁЯПлЁЯОУ

тПня╕П Next

The other schools, where the boxes on your plan are built:

git checkout main
python3 design/demo.py     # one last run, for fun
тЖР Previousdesigns short notify chatFinished! Take the quiz тЖТcheck what stuck

This page is the lesson's README from the lesson-18-designs-files-video-rag branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.