ЁЯПл The SchoolтА║ЁЯУЭ TestingтА║ЁЯзм рдзрдбрд╛ 10 тАФ Coverage рд╡рд┐рд░реБрджреНрдз mutation testing: 100% рдЬреА рдлрдХреНрдд рдХрд╛рд╣реА mutants рдорд╛рд░рддреЗ
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯзм рдзрдбрд╛ 10 тАФ Coverage рд╡рд┐рд░реБрджреНрдз mutation testing: 100% рдЬреА рдлрдХреНрдд рдХрд╛рд╣реА mutants рдорд╛рд░рддреЗ

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 10 ┬╖ рдорд╛рдЧреАрд▓: lesson-09-property-based ┬╖ рдкреБрдвреАрд▓: lesson-11-flaky-tests


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ09, рдЖрдгрд┐ рддреНрдпрд╛рд╢рд┐рд╡рд╛рдп tests рдЪреАрдЪ рдкрд░реАрдХреНрд╖рд╛ рдШреЗрдгреНрдпрд╛рдЪреЗ рджреЛрди рдорд╛рд░реНрдЧ: line coverage (suite рдЪрд╛рд▓рддрд╛рдирд╛ grade() рдЪреНрдпрд╛ рдХреЛрдгрддреНрдпрд╛ lines рдЪрд╛рд▓рд▓реНрдпрд╛) рдЖрдгрд┐ mutation testing (grade() рдордзреНрдпреЗ рдЫреЛрдЯреЗ bugs тАФ mutants тАФ рдкреЗрд░рд╛ рдЖрдгрд┐ suite рддреНрдпрд╛рддрд▓реЗ рдХрд┐рддреА рдкрдХрдбрддреЛ рддреЗ рдореЛрдЬрд╛). 100% line coverage рдЕрд╕рд▓реЗрд▓рд╛ suite 10 рдкреИрдХреА рдлрдХреНрдд 4 mutants рдорд╛рд░рддреЛ. line_coverage(), MUTANTS, mutation_run(), WEAK рдЖрдгрд┐ STRONG exam/lab.py рдордзреНрдпреЗ; mutation() exam/demo.py рдордзреНрдпреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдХрддрд░рд┐рдирд╛рд▓рд╛ рдлрдХреНрдд рд╡рд┐рджреНрдпрд╛рд░реНрдерд┐рдиреА рдирд╡реНрд╣реЗ, рддрд░ рддрд┐рдЪреА рдкреНрд░рд╢реНрдирдкрддреНрд░рд┐рдХрд╛ рдЪрд╛рдВрдЧрд▓реА рдЖрд╣реЗ рдХрд╛ рд╣реЗ рдЬрд╛рдгреВрди рдШреНрдпрд╛рдпрдЪреЗ рдЖрд╣реЗ.

рдкрд╣рд┐рд▓реА рдХрд▓реНрдкрдирд╛: coverage. ЁЯУЛ рдкрд╛рдареНрдпрдкреБрд╕реНрддрдХрд╛рдЪреНрдпрд╛ рдкреНрд░рддреНрдпреЗрдХ рдкрд╛рдирд╛рд▓рд╛ рдХрд┐рдорд╛рди рдПрдХрд╛ рдкреНрд░рд╢реНрдирд╛рдиреЗ рд╕реНрдкрд░реНрд╢ рдХреЗрд▓рд╛ рдХрд╛ рд╣реЗ рддреА рддрдкрд╛рд╕рддреЗ. 7 рдкреНрд░рд╢реНрдирд╛рдВрдЪреА рдХрдордХреБрд╡рдд рдкреНрд░рд╢реНрдирдкрддреНрд░рд┐рдХрд╛ рдкреНрд░рддреНрдпреЗрдХ рдкрд╛рдирд╛рд▓рд╛ рд╕реНрдкрд░реНрд╢ рдХрд░рддреЗ: 100%! ЁЯОЙ

рджреБрд╕рд░реА рдХрд▓реНрдкрдирд╛: mutants. ЁЯзм рдРрд╢реНрд╡рд░реНрдпрд╛ рдЧреБрдкрдЪреВрдк рд╡рд┐рджреНрдпрд╛рд░реНрдерд┐рдиреАрдЪреНрдпрд╛ 10 рдкреНрд░рддреА рдмрдирд╡рддреЗ, рдкреНрд░рддреНрдпреЗрдХреАрдд рдПрдХ рдЫреЛрдЯреА рдЪреВрдХ: рдПрдХреАрд▓рд╛ рд╡рд╛рдЯрддреЗ pass рд╣реЛрдгреНрдпрд╛рд╕рд╛рдареА рдХрд┐рдорд╛рди 35 рдирд╡реНрд╣реЗ рддрд░ 35 рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд рд╣рд╡реЗрдд; рдПрдХ рдЕрд░реНрдзреЗ рдЧреБрдг рдЬреБрдиреНрдпрд╛ рдЪреБрдХреАрдЪреНрдпрд╛ рдкрджреНрдзрддреАрдиреЗ round рдХрд░рддреЗ; рдПрдХ рдЬрд┐рдереЗ "A" рдмрд░реЛрдмрд░ рд╣реЛрддрд╛ рддрд┐рдереЗ "B" рджреЗрддреЗ. рдордЧ рдХрддрд░рд┐рдирд╛рдЪреА рдкреНрд░рд╢реНрдирдкрддреНрд░рд┐рдХрд╛ рдкреНрд░рддреНрдпреЗрдХ рдкреНрд░рддреАрд▓рд╛ рджрд┐рд▓реА рдЬрд╛рддреЗ.

100% coverage рдЕрд╕рд▓реЗрд▓реА рдХрдордХреБрд╡рдд рдкреНрд░рд╢реНрдирдкрддреНрд░рд┐рдХрд╛ 10 рдкреИрдХреА рдлрдХреНрдд 4 рдорд╛рд░рддреЗ. рддрд┐рдиреЗ рдкреНрд░рддреНрдпреЗрдХ рдкрд╛рдирд╛рд▓рд╛ рд╕реНрдкрд░реНрд╢ рдХреЗрд▓рд╛, рдкрдг рдХрдбрд╛рдВрдмрджреНрджрд▓ рдХрдзреАрдЪ рд╡рд┐рдЪрд╛рд░рд▓реЗ рдирд╛рд╣реА. рдзрдбрд╛ 03 рдордзрд▓реНрдпрд╛ boundary рдкреНрд░рд╢реНрдирд╛рдВрд╕рд╣ рдЕрд╕рд▓реЗрд▓реА рдордЬрдмреВрдд рдкреНрд░рд╢реНрдирдкрддреНрд░рд┐рдХрд╛ 9 рдорд╛рд░рддреЗ.

рдЖрдгрд┐ 10рд╡рд╛? рддреНрдпрд╛рдиреЗ "рдХрд┐рдорд╛рди 75" рдЪреЗ "74 рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд" рдХреЗрд▓реЗ. рдкреВрд░реНрдг рд╕рдВрдЦреНрдпрд╛рдВрд╕рд╛рдареА рд╣рд╛ рддреЛрдЪ рдирд┐рдпрдо рдЖрд╣реЗ. рдХреЛрдгрддрд╛рд╣реА рдкреНрд░рд╢реНрди рддреНрдпрд╛рд▓рд╛ рдХрдзреАрдЪ рдкрдХрдбреВ рд╢рдХрдд рдирд╛рд╣реА, рдХрд╛рд░рдг рддреА рдЪреВрдХрдЪ рдирд╛рд╣реА. рддреНрдпрд╛рд▓рд╛рдЪ equivalent mutant рдореНрд╣рдгрддрд╛рдд.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    subgraph weak["ЁЯУД WEAK suite: 7 questions"]
      w["coverage 12/12 = 100%<br/>mutants killed 4/10"]
    end
    subgraph strong["ЁЯУС STRONG suite: 21 questions"]
      s["coverage 12/12 = 100%<br/>mutants killed 9/10"]
    end
    weak -->|"add boundary questions<br/>34.5 ┬╖ 35 ┬╖ 60 ┬╖ 74.5 ┬╖ 75 ┬╖ 100.1"| strong
    strong --> eq["the survivor: s >= 75 тЖТ s > 74<br/>EQUIVALENT for whole numbers"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/testing/lesson-diagrams.html#l10

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг coverage рдореЛрдЬрд╛рдпрд▓рд╛ рд╕реЛрдкреА рдЖрдгрд┐ рдлрд╕рд╡рд╛рдпрд▓рд╛рд╣реА рд╕реЛрдкреА рдЖрд╣реЗ, рдЖрдгрд┐ teams рдЕрд╕реЗ coverage targets рдард░рд╡рддрд╛рдд рдЬреЗ assert рдирд╕рд▓реЗрд▓реНрдпрд╛ tests рд╕реБрджреНрдзрд╛ рдЧрд╛рдареВ рд╢рдХрддрд╛рдд. Mutation testing рддреБрдореНрд╣рд╛рд▓рд╛ рдЦрд░реЛрдЦрд░ рд╣рд╡реЗ рддреЗ рдореЛрдЬрддреЗ: suite рд▓рд╛ bug рд▓рдХреНрд╖рд╛рдд рдпреЗрдИрд▓ рдХрд╛? рдЯрд┐рдХрд▓реЗрд▓реЗ mutants рдореНрд╣рдгрдЬреЗ рдиреЗрдордХреЗ рдХреЛрдгрддреЗ рдкреНрд░рд╢реНрди рд░рд╛рд╣рд┐рд▓реЗ рдЖрд╣реЗрдд рдпрд╛рдЪреА рдпрд╛рджреА тАФ score > 100 тЖТ score > 101 рдЯрд┐рдХрд▓рд╛ рдореНрд╣рдгрдЬреЗ "100.1 рдмрджреНрджрд▓ рдХреЛрдгреАрдЪ рд╡рд┐рдЪрд╛рд░рд▓реЗ рдирд╛рд╣реА".

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

exam/lab.py рдордзреАрд▓ line_coverage(fn, suite) sys.settrace рдиреЗ рдПрдХ trace function рд▓рд╛рд╡рддреЗ, suite рдЪрд╛рд▓рддрд╛рдирд╛ fn рдЪреНрдпрд╛ code object рдордзреАрд▓ рдкреНрд░рддреНрдпреЗрдХ line event рдиреЛрдВрджрд╡рддреЗ, рдЖрдгрд┐ fn рдЪрд╛ source ast рдиреЗ parse рдХрд░реВрди рд╕рд╛рдкрдбрд▓реЗрд▓реНрдпрд╛ statement lines рд╢реА рддреБрд▓рдирд╛ рдХрд░рддреЗ. MUTANTS рдордзреНрдпреЗ 10 (рдЬреБрдирд╛ рдордЬрдХреВрд░, рдирд╡рд╛ рдордЬрдХреВрд░) рдЬреЛрдбреНрдпрд╛ рдЖрд╣реЗрдд; mutant(old, new) grade() рдЪрд╛ source рдкреБрдиреНрд╣рд╛ рд▓рд┐рд╣рд┐рддреЗ рдЖрдгрд┐ рдирд╡реНрдпрд╛ namespace рдордзреНрдпреЗ exec рдХрд░рддреЗ. mutation_run(suite) рдкреНрд░рддреНрдпреЗрдХ mutant рд╡рд░ suite рдЪрд╛рд▓рд╡рддреЗ рдЖрдгрд┐ рдХреЛрдгрддрд╛ case fail рдЭрд╛рд▓рд╛ рдХрд╛ рддреЗ рдиреЛрдВрджрд╡рддреЗ. WEAK рдордзреНрдпреЗ рд╡рд░реНрдЧрд╛рдЪреНрдпрд╛ рдордзрд▓реНрдпрд╛ рднрд╛рдЧрд╛рддрд▓реЗ 7 cases рдЖрд╣реЗрдд; STRONG рддреНрдпрд╛рдд рдзрдбрд╛ 03 рдордзрд▓рд╛ boundary table рдЬреЛрдбрддреЛ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 exam/demo.py mutation
python3 - <<'EOF'
from exam.lab import line_coverage, mutation_run, statement_lines, WEAK, STRONG
from exam.grades import grade
print("statement lines of grade():", statement_lines(grade))
for name, suite in (("one question", [(90, "A")]), ("WEAK", WEAK), ("WEAK + 3 edges", WEAK + [(34.5, "D"), (75, "A"), (100.1, ValueError)]), ("STRONG", STRONG)):
    hit, total = line_coverage(grade, suite)
    res = mutation_run(suite)
    print(f"{name:<15} coverage {hit:>2}/{total} ┬╖ killed {sum(k for *_, k in res):>2}/10 ┬╖ survivors {[n for o, n, k in res if not k]}")
EOF

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

mutation рд╣реЗ рдЫрд╛рдкрддреЗ:

тФАтФА WEAK suite (7 questions): line coverage of grade() 12/12 = 100% ┬╖ mutants killed 4/10
   s >= 75               тЖТ s > 75         SURVIVED
   s >= 35               тЖТ s > 35         SURVIVED
   s >= 60               тЖТ s >= 61        SURVIVED
   round_half_up(score)  тЖТ round(score)   SURVIVED
   return "A"            тЖТ return "B"     killed
тФАтФА STRONG suite (21 questions): line coverage of grade() 12/12 = 100% ┬╖ mutants killed 9/10

рддреБрдордЪрд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

statement lines of grade(): [27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38]
one question    coverage  4/12 ┬╖ killed  1/10 ┬╖ survivors ['s > 75', 's > 35', 's >= 61', 'round(score)', 'return "D"', 'score < 0 and', 's <= 45', 'score > 101', 's > 74']
WEAK            coverage 12/12 ┬╖ killed  4/10 ┬╖ survivors ['s > 75', 's > 35', 's >= 61', 'round(score)', 'score > 101', 's > 74']
WEAK + 3 edges  coverage 12/12 ┬╖ killed  8/10 ┬╖ survivors ['s >= 61', 's > 74']
STRONG          coverage 12/12 ┬╖ killed  9/10 ┬╖ survivors ['s > 74']

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

7 рдкреНрд░рд╢реНрдирд╛рдВрдирдВрддрд░ coverage рдЪреА рдорджрдд рдерд╛рдВрдмрд▓реА: WEAK рдкрд╛рд╕реВрди STRONG рдкрд░реНрдпрдВрдд рддреА 100% рд╡рд░рдЪ рд░рд╛рд╣рд┐рд▓реА, рддрд░ рдорд╛рд░рд▓реЗрд▓реНрдпрд╛рдВрдЪреА рд╕рдВрдЦреНрдпрд╛ 4 рд╡рд░реВрди 9 рдЭрд╛рд▓реА. WEAK suite рдиреЗ рдзрдбрд╛ 01 рдордзреНрдпреЗ ship рдЭрд╛рд▓реЗрд▓рд╛ рдиреЗрдордХрд╛ bug тАФ round_half_up(score) тЖТ round(score) тАФ рдкреВрд░реНрдг coverage рдЕрд╕реВрдирд╣реА рдЯрд┐рдХреВ рджрд┐рд▓рд╛. рдиреАрдЯ рдирд┐рд╡рдбрд▓реЗрд▓реНрдпрд╛ рддреАрди рдХрдбреЗрдЪреНрдпрд╛ рдкреНрд░рд╢реНрдирд╛рдВрдиреА рдЖрдгрдЦреА рдЪрд╛рд░ mutants рдорд╛рд░рд▓реЗ, рдЖрдгрд┐ рдЙрд░рд▓реЗрд▓реНрдпрд╛ рдкреНрд░рддреНрдпреЗрдХ survivor рдиреЗ рдЕрдЬреВрди рдХреЛрдгрддрд╛ рдкреНрд░рд╢реНрди рд░рд╛рд╣рд┐рд▓рд╛ рдЖрд╣реЗ рддреЗ рд╕рд╛рдВрдЧрд┐рддрд▓реЗ (s >= 61 рд╕рд╛рдареА рдиреЗрдордХреНрдпрд╛ 60 рд╡рд░ рдкреНрд░рд╢реНрди рд╣рд╡рд╛). рд╢реЗрд╡рдЯрдЪрд╛, s > 74, equivalent рдЖрд╣реЗ: рддрд┐рдереЗ рдерд╛рдВрдмрд╛.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ project рдордзреНрдпреЗ тАФ coverage.py рдХрд┐рдВрд╡рд╛ pytest-cov plugin рдиреЗ coverage; -m рдЖрдгрд┐ term-missing рдХрдзреАрдЪ рди рдЪрд╛рд▓рд▓реЗрд▓реНрдпрд╛ lines рдЪреА рдпрд╛рджреА рджреЗрддрд╛рдд:

pip install coverage pytest pytest-cov
coverage run --branch -m pytest && coverage report -m
pytest --cov=exam --cov-branch --cov-report=term-missing --cov-fail-under=90

mutmut рдиреЗ mutation testing (Python; рддреЗ рддреБрдордЪрд╛ pytest suite рдкреНрд░рддреНрдпреЗрдХ mutant рд╡рд░ рдЪрд╛рд▓рд╡рддреЗ):

pip install mutmut
mutmut run            # generate mutants and run the tests against each one
mutmut results        # list the survivors
mutmut show <mutant>  # the diff of one survivor

рдЗрддрд░ tools: Java рд╕рд╛рдареА PIT (PITest), JavaScript/TypeScript, C# рдЖрдгрд┐ Scala рд╕рд╛рдареА Stryker, Python рд╕рд╛рдареА cosmic-ray. Google рдиреЗ рд╕рд╛рдВрдЧрд┐рддрд▓реЗ рдЖрд╣реЗ рдХреА рддреЗ рдЯрд┐рдХрд▓реЗрд▓реЗ mutants project-wide score рдореНрд╣рдгреВрди рдирд╡реНрд╣реЗ, рддрд░ code review рдордзреНрдпреЗ comments рдореНрд╣рдгреВрди, рдлрдХреНрдд рдмрджрд▓рд▓реЗрд▓реНрдпрд╛ lines рд╡рд░ рджрд╛рдЦрд╡рддрд╛рдд.

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: test рди рдЭрд╛рд▓реЗрд▓реНрдпрд╛ code рд╕рд╛рдареА coverage рдПрдХ рд╕реНрд╡рд╕реНрдд рдзреЛрдХреНрдпрд╛рдЪреА рдШрдВрдЯрд╛ рдореНрд╣рдгреВрди рдареЗрд╡рд╛, рдЖрдгрд┐ рдЬрд┐рдереЗ рдЪреБрдХреАрдЪреЗ рдЙрддреНрддрд░ рдорд╣рд╛рдЧ рдкрдбрддреЗ рдЕрд╢рд╛ рдереЛрдбреНрдпрд╛ modules рд╡рд░ mutation testing рдЪрд╛рд▓рд╡рд╛ тАФ grading, billing, permissions. рдкреНрд░рддреНрдпреЗрдХ survivor рдореНрд╣рдгрдЬреЗ рдПрдХрддрд░ рдирд╕рд▓реЗрд▓реА test рдХрд┐рдВрд╡рд╛ equivalent mutant; рджреЛрдиреНрд╣реА рдорд╛рд╣реАрдд рдЕрд╕рдгреЗ рдлрд╛рдпрджреНрдпрд╛рдЪреЗ рдЖрд╣реЗ.

тПня╕П рдкреБрдвреЗ

рдЪрд╛рдВрдЧрд▓рд╛ suite рд╕реБрджреНрдзрд╛ рджреБрд╕рд▒реНрдпрд╛ рдкреНрд░рдХрд╛рд░реЗ рдЦреЛрдЯреЗ рдмреЛрд▓реВ рд╢рдХрддреЛ: code рдордзреНрдпреЗ рдХрд╛рд╣реАрд╣реА рдмрджрд▓ рди рдХрд░рддрд╛ рд╕реЛрдорд╡рд╛рд░реА pass рдЖрдгрд┐ рдордВрдЧрд│рд╡рд╛рд░реА fail рд╣реЛрдгрд╛рд░реА test. рднрд┐рдВрддреАрд╡рд░рдЪреЗ рдбреБрдЧрдбреБрдЧрдгрд╛рд░реЗ рдШрдбреНрдпрд╛рд│ тАФ flaky tests.

git checkout lesson-11-flaky-tests

ЁЯзм Lesson 10 тАФ Coverage vs mutation testing: 100% that kills only some mutants

ЁЯУН You are here: Lesson 10 of 12 ┬╖ Previous: lesson-09-property-based ┬╖ Next: lesson-11-flaky-tests


ЁЯУж What's in this branch

Lessons 01тАУ09, plus two ways to grade the tests themselves: line coverage (which lines of grade() ran while the suite ran) and mutation testing (plant small bugs тАФ mutants тАФ in grade() and count how many the suite catches). A suite with 100% line coverage kills only 4 of 10 mutants. line_coverage(), MUTANTS, mutation_run(), WEAK and STRONG in exam/lab.py; mutation() in exam/demo.py.

ЁЯзТ Explain like I'm 5

Katrina wants to know if her exam paper is good, not just the pupil.

First idea: coverage. ЁЯУЛ She checks that every page of the textbook was touched by at least one question. The weak paper, with 7 questions, touches every page: 100%! ЁЯОЙ

Second idea: mutants. ЁЯзм Aishwarya secretly makes 10 copies of the pupil, each with one small mistake: one thinks a pass needs more than 35 instead of at least 35; one rounds halves the old wrong way; one gives "B" where "A" was right. Then Katrina's paper is given to each copy.

The weak paper, with its 100% coverage, kills only 4 of the 10. It touched every page but never asked about the edges. The strong paper, with the boundary questions from lesson 03, kills 9.

And the 10th? It changed "at least 75" to "more than 74". For whole numbers, that is the same rule. No question can ever catch it, because it is not a mistake. That is an equivalent mutant.

ЁЯЧ║я╕П Diagram

flowchart LR
    subgraph weak["ЁЯУД WEAK suite: 7 questions"]
      w["coverage 12/12 = 100%<br/>mutants killed 4/10"]
    end
    subgraph strong["ЁЯУС STRONG suite: 21 questions"]
      s["coverage 12/12 = 100%<br/>mutants killed 9/10"]
    end
    weak -->|"add boundary questions<br/>34.5 ┬╖ 35 ┬╖ 60 ┬╖ 74.5 ┬╖ 75 ┬╖ 100.1"| strong
    strong --> eq["the survivor: s >= 75 тЖТ s > 74<br/>EQUIVALENT for whole numbers"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/testing/lesson-diagrams.html#l10

тЭУ What

ЁЯдФ Why

Because coverage is easy to measure and easy to game, and teams set coverage targets that tests without asserts can meet. Mutation testing measures what you actually want: would the suite notice a bug? The survivors are a to-do list of exactly which questions are missing тАФ score > 100 тЖТ score > 101 surviving says "nobody asked about 100.1".

ЁЯФз How (in this repo)

line_coverage(fn, suite) in exam/lab.py sets a trace function with sys.settrace, records every line event inside fn's code object while the suite runs, and compares with the statement lines found by parsing fn's source with ast. MUTANTS lists 10 (old text, new text) pairs; mutant(old, new) rewrites the source of grade() and execs it into a fresh namespace. mutation_run(suite) runs the suite against each mutant and records whether any case failed. WEAK has 7 middle-of-class cases; STRONG adds the boundary table from lesson 03.

ЁЯзк Try it

python3 exam/demo.py mutation
python3 - <<'EOF'
from exam.lab import line_coverage, mutation_run, statement_lines, WEAK, STRONG
from exam.grades import grade
print("statement lines of grade():", statement_lines(grade))
for name, suite in (("one question", [(90, "A")]), ("WEAK", WEAK), ("WEAK + 3 edges", WEAK + [(34.5, "D"), (75, "A"), (100.1, ValueError)]), ("STRONG", STRONG)):
    hit, total = line_coverage(grade, suite)
    res = mutation_run(suite)
    print(f"{name:<15} coverage {hit:>2}/{total} ┬╖ killed {sum(k for *_, k in res):>2}/10 ┬╖ survivors {[n for o, n, k in res if not k]}")
EOF

тЬЕ Verify тАФ what you should see

mutation prints:

тФАтФА WEAK suite (7 questions): line coverage of grade() 12/12 = 100% ┬╖ mutants killed 4/10
   s >= 75               тЖТ s > 75         SURVIVED
   s >= 35               тЖТ s > 35         SURVIVED
   s >= 60               тЖТ s >= 61        SURVIVED
   round_half_up(score)  тЖТ round(score)   SURVIVED
   return "A"            тЖТ return "B"     killed
тФАтФА STRONG suite (21 questions): line coverage of grade() 12/12 = 100% ┬╖ mutants killed 9/10

Your snippet prints:

statement lines of grade(): [27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38]
one question    coverage  4/12 ┬╖ killed  1/10 ┬╖ survivors ['s > 75', 's > 35', 's >= 61', 'round(score)', 'return "D"', 'score < 0 and', 's <= 45', 'score > 101', 's > 74']
WEAK            coverage 12/12 ┬╖ killed  4/10 ┬╖ survivors ['s > 75', 's > 35', 's >= 61', 'round(score)', 'score > 101', 's > 74']
WEAK + 3 edges  coverage 12/12 ┬╖ killed  8/10 ┬╖ survivors ['s >= 61', 's > 74']
STRONG          coverage 12/12 ┬╖ killed  9/10 ┬╖ survivors ['s > 74']

ЁЯПБ What you just proved

Coverage stopped helping at 7 questions: from WEAK to STRONG it stayed at 100% while the kill count went from 4 to 9. The WEAK suite let the exact bug that shipped in lesson 01 тАФ round_half_up(score) тЖТ round(score) тАФ survive with full coverage. Three well-chosen edge questions killed four more mutants, and each remaining survivor named the question still missing (s >= 61 needs a question at exactly 60). The last one, s > 74, is equivalent: stop there.

тЪая╕П Common mistakes

ЁЯПн In production

On a real project тАФ coverage with coverage.py or the pytest-cov plugin; -m and term-missing list the lines that never ran:

pip install coverage pytest pytest-cov
coverage run --branch -m pytest && coverage report -m
pytest --cov=exam --cov-branch --cov-report=term-missing --cov-fail-under=90

Mutation testing with mutmut (Python; it runs your pytest suite against each mutant):

pip install mutmut
mutmut run            # generate mutants and run the tests against each one
mutmut results        # list the survivors
mutmut show <mutant>  # the diff of one survivor

Other tools: PIT (PITest) for Java, Stryker for JavaScript/TypeScript, C# and Scala, cosmic-ray for Python. Google has described showing surviving mutants as comments in code review, on changed lines only, rather than a project-wide score.

ЁЯПн Why this matters in production: keep coverage as a cheap alarm for untested code, and run mutation testing on the few modules where a wrong answer is expensive тАФ grading, billing, permissions. Every survivor is either a missing test or an equivalent mutant; both are worth knowing.

тПня╕П Next

A good suite can still lie in another way: a test that passes on Monday and fails on Tuesday with no code change. The wobbly clock on the wall тАФ flaky tests.

git checkout lesson-11-flaky-tests
тЖР Previousproperty basedNext тЖТflaky tests

This page is the lesson's README from the lesson-10-coverage-mutation branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.