ЁЯПл The SchoolтА║ЁЯПОя╕П PerformanceтА║ЁЯз╡ рдзрдбрд╛ 08 тАФ Concurrency рдЖрдгрд┐ GIL: рдПрдХ рдмреЕрдЯрди, рдЕрдиреЗрдХ рдзрд╛рд╡рдкрдЯреВ
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯз╡ рдзрдбрд╛ 08 тАФ Concurrency рдЖрдгрд┐ GIL: рдПрдХ рдмреЕрдЯрди, рдЕрдиреЗрдХ рдзрд╛рд╡рдкрдЯреВ

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 08 ┬╖ рдорд╛рдЧреЗ: lesson-07-io-batching ┬╖ рдкреБрдвреЗ: lesson-09-queueing


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ07, рдЖрдгрд┐ concurrency: рдПрдХрд╛рдЪ рд╡реЗрд│реА рдЕрдиреЗрдХ рдХрд╛рдореЗ рдХрд░рдгреЗ. рдкрд╣рд┐рд▓рд╛ рдкреНрд░рд╢реНрди рдиреЗрд╣рдореА рд╣рд╛рдЪ рдЕрд╕рддреЛ: рдХрд╛рдо рдХрд╢рд╛рдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рдд рдЖрд╣реЗ. I/O-bound рдХрд╛рдо рдЙрддреНрддрд░рд╛рдВрдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рддреЗ, рдЖрдгрд┐ threads рдХрд┐рдВрд╡рд╛ asyncio рд╣реА рд╡рд╛рдЯ рдкрд╛рд╣рдгреЗ рдПрдХрдореЗрдХрд╛рдВрд╡рд░ рдЪрдврд╡реВ (overlap) рд╢рдХрддрд╛рдд. CPU-bound рдХрд╛рдорд╛рд▓рд╛ processor рд▓рд╛рдЧрддреЛ, рдЖрдгрд┐ standard CPython рдордзреНрдпреЗ рдПрдХрд╛ рд╡реЗрд│реА рдПрдХрдЪ thread Python code рдЪрд╛рд▓рд╡рддреЛ тАФ рд╣рд╛рдЪ GIL. рддреНрдпрд╛рд╕рд╛рдареА рддреБрдореНрд╣рд╛рд▓рд╛ processes (рдХрд┐рдВрд╡рд╛ free-threaded build) рд▓рд╛рдЧрддрд╛рдд. perf/demo.py рдордзреАрд▓ concurrency() рдЖрдгрд┐ perf/sim.py рдордзреАрд▓ schedule тАФ рд╣реЗ tick-by-tick рдЪрд╛рд▓рдгрд╛рд░реЗ рд╢рд┐рдХрд╡рдгреНрдпрд╛рд╕рд╛рдареАрдЪреЗ model рдЖрд╣реЗ, рдЦрд░реЗ threads рдирд╡реНрд╣реЗрдд.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рд░рд┐рд▓реЗ рд╕рдВрдШрд╛рдд 8 рдзрд╛рд╡рдкрдЯреВ рдЖрд╣реЗрдд рдЖрдгрд┐ рдПрдХрдЪ рдмреЕрдЯрди. ЁЯев рдпрд╛ track рд╡рд░ рдзрд╛рд╡рдкрдЯреВ рдлрдХреНрдд рдмреЕрдЯрди рд╣рд╛рддрд╛рдд рдЕрд╕рддрд╛рдирд╛рдЪ рдзрд╛рд╡реВ рд╢рдХрддреЗ.

Sprints рд╡реЗрдЧрд╡рд╛рди рдХрд╕реЗ рдХрд░рд╛рдпрдЪреЗ? рдЕрдзрд┐рдХ tracks рд╡рд╛рдкрд░рд╛, рдкреНрд░рддреНрдпреЗрдХрд╛рд╡рд░ рд╕реНрд╡рддрдГрдЪреЗ рдмреЕрдЯрди (рдЕрдиреЗрдХ cores рд╡рд░ processes). рдХрд┐рдВрд╡рд╛ рдЕрд╕рд╛ рдирд╡рд╛ track рдЬрд┐рдереЗ рдкреНрд░рддреНрдпреЗрдХ рдзрд╛рд╡рдкрдЯреВ рдПрдХрд╛рдЪ рд╡реЗрд│реА рдзрд╛рд╡реВ рд╢рдХрддреЗ (free-threaded build). рдорд╛рддреНрд░ рджреБрд╕рд▒реНрдпрд╛ tracks рд╡рд░ рдкреЛрд╣реЛрдЪрд╛рдпрд▓рд╛ рдЖрдзреА рдереЛрдбрд╛ рд╡реЗрд│ рд▓рд╛рдЧрддреЛ.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    io["ЁЯЪ░ I/O-bound ├Ч 8<br/>5 ms Python + 100 ms waiting"] --> seq1["one after another: 840 ms"]
    io --> thr1["8 threads (GIL): 140 ms"]
    io --> as1["asyncio: 140 ms"]
    cpu["ЁЯПГ CPU-bound ├Ч 8<br/>100 ms Python"] --> seq2["one after another: 800 ms"]
    cpu --> thr2["8 threads (GIL): 800 ms"]
    cpu --> pr["4 processes: 250 ms"]
    cpu --> ft["free-threaded, 4 cores: 200 ms"]

ЁЯЧ║я╕П рд░реЗрдЦрд╛рдЯрд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/performance/lesson-diagrams.html#l08

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг "threads рд╡рд╛рдврд╡рд╛" рд╣рд╛ performance рдЪрд╛ рд╕рд░реНрд╡рд╛рдд рд╕рд╛рдорд╛рдиреНрдп рд╕рд▓реНрд▓рд╛ рдЖрд╣реЗ рдЖрдгрд┐ рддреЛ рдЕрдиреЗрдХрджрд╛ рдЪреБрдХреАрдЪрд╛ рдЕрд╕рддреЛ. Web scraper рд╕рд╛рдареА threads рдХрд┐рдВрд╡рд╛ asyncio рдореЛрдареА speed-up рджреЗрдК рд╢рдХрддрд╛рдд. рд╢реБрджреНрдз Python рдордзрд▓реНрдпрд╛ рдЖрдХрдбреЗрдореЛрдбреАрдЪреНрдпрд╛ loop рд╕рд╛рдареА, standard build рдордзрд▓реЗ threads рдХрд╛рд╣реАрдЪ speed-up рджреЗрдд рдирд╛рд╣реАрдд, рдЖрдгрд┐ processes рдкреНрд░рддреНрдпреЗрдХреА рдЬрд╛рд╕реНрддреАрдд рдЬрд╛рд╕реНрдд рдПрдХрд╛ core рдЗрддрдХрд╛ рдлрд╛рдпрджрд╛ рджреЗрддрд╛рдд тАФ start-up рдЖрдгрд┐ data copying рд╡рдЬрд╛ рдХрд░реВрди. рд╡рд╛рдЯ рдкрд╛рд╣рдгреНрдпрд╛рдЪреНрдпрд╛ рдкреНрд░рдХрд╛рд░рд╛рдиреБрд╕рд╛рд░ рдирд┐рд╡рдбрд▓реНрдпрд╛рдиреЗ рдЦреВрдк рд╡рд╛рдпрд╛ рдЬрд╛рдгрд╛рд░реЗ рдХрд╛рдо рд╡рд╛рдЪрддреЗ.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

perf/sim.py рдордзреАрд▓ schedule(tasks, workers, cpu_slots, start_ms) рдПрдХ tick-by-tick model рдЪрд╛рд▓рд╡рддреЗ: рдкреНрд░рддреНрдпреЗрдХ task рдореНрд╣рдгрдЬреЗ phases рдЪреА рдпрд╛рджреА, ("cpu", ms) рдХрд┐рдВрд╡рд╛ ("io", ms). рдПрдХрд╛ рд╡реЗрд│реА рдЬрд╛рд╕реНрддреАрдд рдЬрд╛рд╕реНрдд workers tasks рдЪрд╛рд▓реВ рдЕрд╕рддрд╛рдд; рд╡рд╛рдЯ рдкрд╛рд╣рдгреНрдпрд╛рд▓рд╛ (io) CPU рд▓рд╛рдЧрдд рдирд╛рд╣реА; рдПрдХрд╛рдЪ millisecond рдордзреНрдпреЗ рдлрдХреНрдд cpu_slots tasks рдЧрдгрдирд╛ рдХрд░реВ рд╢рдХрддрд╛рдд тАФ GIL рд╕рд╣ threads рдХрд┐рдВрд╡рд╛ asyncio рд╕рд╛рдареА 1, processes рдХрд┐рдВрд╡рд╛ free-threaded build рд╕рд╛рдареА cores рдЪреА рд╕рдВрдЦреНрдпрд╛ тАФ round-robin рдкрджреНрдзрддреАрдиреЗ рд╡рд╛рдЯреВрди. рдкреНрд░рддреНрдпреЗрдХ process рдПрдХрджрд╛рдЪ start_ms рдореЛрдЬрддреЗ. рдЦрд░реЗ CPython рдкреНрд░рддреНрдпреЗрдХ millisecond рд▓рд╛ рдирд╡реНрд╣реЗ рддрд░ рджрд░ рдХрд╛рд╣реА milliseconds рд▓рд╛ switch рдХрд░рддреЗ, рдЖрдгрд┐ рдЦрд▒реНрдпрд╛ I/O рд╡ start-up рдЪреНрдпрд╛ рд╡реЗрд│рд╛ рдмрджрд▓рдд рдЕрд╕рддрд╛рдд; рдкреБрдвреЗ рдЯрд┐рдХрддреЛ рддреЛ рдирд┐рдХрд╛рд▓рд╛рдВрдЪрд╛ рдЖрдХрд╛рд░.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 perf/demo.py concurrency
python3 - <<'EOF'
import sys; sys.path.insert(0, "perf"); from sim import schedule
mixed = [[("cpu", 20), ("io", 80)]] * 8
print("8 mixed tasks (20 ms Python + 80 ms waiting):")
for label, w, slots, start in (("one after another", 1, 1, 0), ("8 threads, GIL", 8, 1, 0), ("4 processes", 4, 4, 50), ("8 processes", 8, 4, 50)):
    print(f"  {label:<18} {schedule(mixed, w, slots, start):>4} ms")
for n in (2, 4, 8, 16):
    print(f"CPU-bound ├Ч 8 on {n:>2} threads with the GIL: {schedule([[('cpu', 100)]] * 8, n, 1)} ms")
EOF
python3 - <<'EOF'
import time, threading
def wait(): time.sleep(0.1)
t0 = time.perf_counter(); [wait() for _ in range(8)]; seq = time.perf_counter() - t0
ts = [threading.Thread(target=wait) for _ in range(8)]
t0 = time.perf_counter(); [t.start() for t in ts]; [t.join() for t in ts]; thr = time.perf_counter() - t0
print(f"real: 8 sleeps of 0.1 s тЖТ one after another {seq:.2f} s ┬╖ 8 threads {thr:.2f} s   (your numbers will differ)")
EOF

рд╢реЗрд╡рдЯрдЪрд╛ snippet рдЦрд░реЗ threads рдЖрдгрд┐ рдЦрд░реЗ рдШрдбреНрдпрд╛рд│ рд╡рд╛рдкрд░рддреЛ тАФ рд╕рд╛рдзрд╛рд░рдг 0.82 s рд╡рд┐рд░реБрджреНрдз 0.11 s. рддреБрдордЪреЗ рдЖрдХрдбреЗ рдереЛрдбреЗ рд╡реЗрдЧрд│реЗ рдЕрд╕рддреАрд▓; 8├Ч рдЪрд╛ рдЖрдХрд╛рд░ рдмрджрд▓рдгрд╛рд░ рдирд╛рд╣реА, рдХрд╛рд░рдг sleep GIL рд╕реЛрдбрддреЛ.

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

concurrency рд╣реЗ рдЫрд╛рдкрддреЗ:

тФАтФА 8 tasks ┬╖ I/O-bound: 5 ms of Python + 100 ms waiting for a reply ┬╖ CPU-bound: 100 ms of Python (a teaching model)
                                             I/O-bound   CPU-bound
   one after another                           840 ms      800 ms
   8 threads, one baton (the GIL)              140 ms      800 ms
   asyncio, 1 thread, 8 tasks                  140 ms      800 ms
   4 processes on 4 cores (+50 ms start)       260 ms      250 ms
   8 threads, free-threaded build, 4 cores     110 ms      200 ms

рддреБрдордЪрд╛ рдкрд╣рд┐рд▓рд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

8 mixed tasks (20 ms Python + 80 ms waiting):
  one after another   800 ms
  8 threads, GIL      240 ms
  4 processes         250 ms
  8 processes         170 ms
CPU-bound ├Ч 8 on  2 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on  4 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on  8 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on 16 threads with the GIL: 800 ms

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

GIL рд╕рд╣, CPU рдХрд╛рдорд╛рд▓рд╛ 2 threads рдЕрд╕реЛрдд рдХреА 16, 800 ms рд▓рд╛рдЧрд▓реЗ тАФ Python рдХрд╛рдо рдлрдХреНрдд рдЖрд│реАрдкрд╛рд│реАрдиреЗ рд╣реЛрддреЗ. рд╡рд╛рдЯ рдкрд╛рд╣рдгреЗ рдкреВрд░реНрдгрдкрдгреЗ overlap рдЭрд╛рд▓реЗ: 840 тЖТ 140 ms. Processes рдиреА CPU рдЪреА рд╢рд░реНрдпрдд рдЬрд┐рдВрдХрд▓реА (250 ms) рдкрдг I/O рдЪреА рд╢рд░реНрдпрдд рд╣рд░рд▓реА (260 ms рд╡рд┐рд░реБрджреНрдз 140 ms), рдХрд╛рд░рдг рддреНрдпрд╛рдВрдЪрд╛ start-up рдЦрд░реНрдЪ рдЖрдгрд┐ рдлрдХреНрдд 4 workers. рдорд┐рд╢реНрд░ рдХрд╛рдорд╛рдд, GIL рдореБрд│реЗ 8 ├Ч 20 ms = 160 ms Python рдПрдХрд╛рдЪ рд░рд╛рдВрдЧреЗрдд рд░рд╛рд╣рд┐рд▓реЗ, рдЕрдзрд┐рдХ рдПрдХ 80 ms рдЪреА рдкреНрд░рддреАрдХреНрд╖рд╛: 240 ms. рддреБрдордЪреЗ рдХрд╛рдо рдХрд╢рд╛рдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рддреЗ рддреЗ рдУрд│рдЦрд╛, рдордЧ рдирд┐рд╡рдбрд╛.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ machine рд╡рд░ тАФ standard library рддрд┐рдиреНрд╣реА рдкреНрд░рдХрд╛рд░ рд╣рд╛рддрд╛рд│рддреЗ:

from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import asyncio

with ThreadPoolExecutor(max_workers=16) as pool:          # I/O-bound: overlap the waits
    pages = list(pool.map(fetch_page, urls))

with ProcessPoolExecutor() as pool:                       # CPU-bound: one process per core
    scores = list(pool.map(score_heat, heats, chunksize=64))

async def main():                                         # thousands of waiting connections
    return await asyncio.gather(*(fetch_async(u) for u in urls))
asyncio.run(main())

рддреБрдордЪреНрдпрд╛рдХрдбреЗ рдХреЛрдгрддрд╛ interpreter рдЖрд╣реЗ рддреЗ рддрдкрд╛рд╕рд╛, рдЖрдгрд┐ free-threaded build (3.13+) рд╡рд╛рдкрд░реВрди рдкрд╛рд╣рд╛:

python3 -c "import sys; print(sys.version); print('GIL enabled:', getattr(sys, '_is_gil_enabled', lambda: True)())"
python3.13t -X gil=0 my_cpu_job.py        # a free-threaded build, if installed; PYTHON_GIL=0 does the same

Web service рд╕рд╛рдареА server рдЪреЗ worker model рд╣реЗ рддреБрдордЪреНрдпрд╛рд╕рд╛рдареА рдард░рд╡рддреЗ тАФ рдЙрджрд╛рд╣рд░рдгрд╛рд░реНрде рдЕрдиреЗрдХ worker processes (parallel CPU) рдЕрд╕рд▓реЗрд▓реЗ Gunicorn, рдкреНрд░рддреНрдпреЗрдХрд╛рдд threads рдХрд┐рдВрд╡рд╛ async loop (I/O overlap рдХрд░рдгрд╛рд░реЗ):

gunicorn app:app --workers 4 --threads 8

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: process рдЪреНрдпрд╛ CPU time рдЪреА рддреНрдпрд╛рдЪреНрдпрд╛ wall time рд╢реА рддреБрд▓рдирд╛ рдХрд░рд╛. рдЬрд░ рдПрдХрд╛ core рд╡рд░ CPU тЙИ wall рдЕрд╕реЗрд▓, рддрд░ рддреЗ CPU-bound рдЖрд╣реЗ: processes, рдЪрд╛рдВрдЧрд▓реЗ algorithms, рдХрд┐рдВрд╡рд╛ native code рдЪрд╛ рд╡рд┐рдЪрд╛рд░ рдХрд░рд╛. рдЬрд░ CPU тЙк wall рдЕрд╕реЗрд▓, рддрд░ рддреЗ рд╡рд╛рдЯ рдкрд╛рд╣рдд рдЖрд╣реЗ: рдкреНрд░рддреАрдХреНрд╖рд╛ overlap рдХрд░рд╛, рддреНрдпрд╛рдВрдЪреЗ batching рдХрд░рд╛ (рдзрдбрд╛ 07), рдХрд┐рдВрд╡рд╛ cache рдХрд░рд╛ (рдзрдбрд╛ 06).

тПня╕П рдкреБрдвреЗ

рдПрдХрд╛рдЪ рдкрд╛рдгреНрдпрд╛рдЪреНрдпрд╛ рдЯреЗрдмрд▓рд╛рдЪрд╛ рд╡рд╛рдкрд░ рдЕрдзрд┐рдХ рдзрд╛рд╡рдкрдЯреВ рдХрд░реВ рд▓рд╛рдЧрд▓реНрдпрд╛ рдХреА рдПрдХ рдирд╡реА рд╕рдорд╕реНрдпрд╛ рджрд┐рд╕рддреЗ: рд░рд╛рдВрдЧ. рдкреБрдвреЗ: utilisation рд╡рд┐рд░реБрджреНрдз рдкреНрд░рддреАрдХреНрд╖реЗрдЪрд╛ рд╡реЗрд│, Little рдЪрд╛ рдирд┐рдпрдо, рдЖрдгрд┐ Amdahl рдЪрд╛ рд╕рд░реНрд╡рд╛рдд рд╣рд│реВ рдзрд╛рд╡рдкрдЯреВ.

git checkout lesson-09-queueing

ЁЯз╡ Lesson 08 тАФ Concurrency & the GIL: one baton, many runners

ЁЯУН You are here: Lesson 08 of 12 ┬╖ Previous: lesson-07-io-batching ┬╖ Next: lesson-09-queueing


ЁЯУж What's in this branch

Lessons 01тАУ07, plus concurrency: doing several things at once. The first question is always what the work is waiting for. I/O-bound work waits for replies, and threads or asyncio can overlap the waiting. CPU-bound work needs the processor, and in standard CPython only one thread runs Python code at a time тАФ the GIL. For that you need processes (or a free-threaded build). concurrency() in perf/demo.py and schedule in perf/sim.py тАФ a tick-by-tick teaching model, not real threads.

ЁЯзТ Explain like I'm 5

The relay team has 8 runners and one baton. ЁЯев On this track, a runner may only run while holding the baton.

How do you make sprints faster? Use more tracks, each with its own baton (processes on several cores). Or a new kind of track where every runner can run at once (the free-threaded build). Getting to the other tracks takes a little time first, though.

ЁЯЧ║я╕П Diagram

flowchart LR
    io["ЁЯЪ░ I/O-bound ├Ч 8<br/>5 ms Python + 100 ms waiting"] --> seq1["one after another: 840 ms"]
    io --> thr1["8 threads (GIL): 140 ms"]
    io --> as1["asyncio: 140 ms"]
    cpu["ЁЯПГ CPU-bound ├Ч 8<br/>100 ms Python"] --> seq2["one after another: 800 ms"]
    cpu --> thr2["8 threads (GIL): 800 ms"]
    cpu --> pr["4 processes: 250 ms"]
    cpu --> ft["free-threaded, 4 cores: 200 ms"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/performance/lesson-diagrams.html#l08

тЭУ What

ЁЯдФ Why

Because "add threads" is the most common performance advice and it is often wrong. For a web scraper, threads or asyncio can give a big speed-up. For a number-crunching loop in pure Python, threads in the standard build give none, and processes give up to one core's worth each тАФ minus start-up and data copying. Choosing by the kind of waiting saves a lot of wasted work.

ЁЯФз How (in this repo)

schedule(tasks, workers, cpu_slots, start_ms) in perf/sim.py runs a tick-by-tick model: each task is a list of phases, ("cpu", ms) or ("io", ms). Up to workers tasks are in progress at once; waiting (io) needs no CPU; only cpu_slots tasks may compute in the same millisecond тАФ 1 for threads or asyncio with the GIL, the number of cores for processes or a free-threaded build тАФ shared round-robin. Each process pays start_ms once. Real CPython switches every few milliseconds, not every one, and real I/O and start-up times vary; the shape of the results is what carries over.

ЁЯзк Try it

python3 perf/demo.py concurrency
python3 - <<'EOF'
import sys; sys.path.insert(0, "perf"); from sim import schedule
mixed = [[("cpu", 20), ("io", 80)]] * 8
print("8 mixed tasks (20 ms Python + 80 ms waiting):")
for label, w, slots, start in (("one after another", 1, 1, 0), ("8 threads, GIL", 8, 1, 0), ("4 processes", 4, 4, 50), ("8 processes", 8, 4, 50)):
    print(f"  {label:<18} {schedule(mixed, w, slots, start):>4} ms")
for n in (2, 4, 8, 16):
    print(f"CPU-bound ├Ч 8 on {n:>2} threads with the GIL: {schedule([[('cpu', 100)]] * 8, n, 1)} ms")
EOF
python3 - <<'EOF'
import time, threading
def wait(): time.sleep(0.1)
t0 = time.perf_counter(); [wait() for _ in range(8)]; seq = time.perf_counter() - t0
ts = [threading.Thread(target=wait) for _ in range(8)]
t0 = time.perf_counter(); [t.start() for t in ts]; [t.join() for t in ts]; thr = time.perf_counter() - t0
print(f"real: 8 sleeps of 0.1 s тЖТ one after another {seq:.2f} s ┬╖ 8 threads {thr:.2f} s   (your numbers will differ)")
EOF

The last snippet uses real threads and a real clock тАФ something like 0.82 s vs 0.11 s. Your numbers will differ a little; the 8├Ч shape will not, because sleep releases the GIL.

тЬЕ Verify тАФ what you should see

concurrency prints:

тФАтФА 8 tasks ┬╖ I/O-bound: 5 ms of Python + 100 ms waiting for a reply ┬╖ CPU-bound: 100 ms of Python (a teaching model)
                                             I/O-bound   CPU-bound
   one after another                           840 ms      800 ms
   8 threads, one baton (the GIL)              140 ms      800 ms
   asyncio, 1 thread, 8 tasks                  140 ms      800 ms
   4 processes on 4 cores (+50 ms start)       260 ms      250 ms
   8 threads, free-threaded build, 4 cores     110 ms      200 ms

Your first snippet prints:

8 mixed tasks (20 ms Python + 80 ms waiting):
  one after another   800 ms
  8 threads, GIL      240 ms
  4 processes         250 ms
  8 processes         170 ms
CPU-bound ├Ч 8 on  2 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on  4 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on  8 threads with the GIL: 800 ms
CPU-bound ├Ч 8 on 16 threads with the GIL: 800 ms

ЁЯПБ What you just proved

With the GIL, CPU work took 800 ms whether there were 2 threads or 16 тАФ the Python work simply takes turns. Waiting overlapped perfectly: 840 тЖТ 140 ms. Processes won the CPU race (250 ms) but lost the I/O race (260 ms vs 140 ms) because of their start-up cost and only 4 workers. For mixed work, the GIL left 8 ├Ч 20 ms = 160 ms of Python in a single line, plus one 80 ms wait: 240 ms. Know what your work waits for, then choose.

тЪая╕П Common mistakes

ЁЯПн In production

On a real machine тАФ the standard library covers the three shapes:

from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import asyncio

with ThreadPoolExecutor(max_workers=16) as pool:          # I/O-bound: overlap the waits
    pages = list(pool.map(fetch_page, urls))

with ProcessPoolExecutor() as pool:                       # CPU-bound: one process per core
    scores = list(pool.map(score_heat, heats, chunksize=64))

async def main():                                         # thousands of waiting connections
    return await asyncio.gather(*(fetch_async(u) for u in urls))
asyncio.run(main())

Check which interpreter you have, and try a free-threaded build (3.13+):

python3 -c "import sys; print(sys.version); print('GIL enabled:', getattr(sys, '_is_gil_enabled', lambda: True)())"
python3.13t -X gil=0 my_cpu_job.py        # a free-threaded build, if installed; PYTHON_GIL=0 does the same

For a web service, the server's worker model decides this for you тАФ for example Gunicorn with several worker processes (CPU in parallel), each with threads or an async loop (overlapping I/O):

gunicorn app:app --workers 4 --threads 8

ЁЯПн Why this matters in production: compare a process's CPU time with its wall time. If CPU тЙИ wall on one core, it is CPU-bound: think processes, better algorithms, or native code. If CPU тЙк wall, it is waiting: overlap the waits, batch them (lesson 07), or cache (lesson 06).

тПня╕П Next

With more runners sharing one water table, a new problem appears: the queue. Next: utilisation vs waiting time, Little's law, and Amdahl's slowest runner.

git checkout lesson-09-queueing
тЖР Previousio batchingNext тЖТqueueing

This page is the lesson's README from the lesson-08-concurrency branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.