code जे करायला हवे तेच करतो हे कसे तपासायचे, हे शाळेच्या परीक्षा हॉलच्या रूपात शिकवले आहे: test म्हणजे परीक्षेचे प्रश्न, fixture म्हणजे तयार ठेवलेल्या उत्तरपत्रिका, test double म्हणजे बदली परीक्षक, आणि flaky test म्हणजे भिंतीवरचे डुगडुगणारे घड्याळ. प्रत्येक धडा ही शाळेतील एक गोष्ट आहे, सोबत एक काढलेली आकृती आणि एक lab — आणि परीक्षा हॉल repo मध्येच आहे: शोधण्यासाठी खरे bug असलेले grading module (exam/grades.py) आणि testing tools शुद्ध Python मधील deterministic models म्हणून (exam/lab.py, शून्य dependencies).
# the 60-second wow — one grading module, twelve ways to test it:
git clone https://github.com/BaluRaut/learn-testing-school.git && cd learn-testing-school
python3 exam/demo.py # 12 lessons: a rounding bug, a lying stub, 10 mutants, a wobbly clock
python3 exam/test_exam.py # 12 checks across the lessons
तुम्हाला काय दिसायला हवे (छाटलेले — यश असे दिसते):
═══ why ═══
── Dipika's grade_v1 sits 5 easy exam questions: 5/5 correct — and it still has a bug
the whole school sits a paper out of 200 (1200 students): 15 report cards get the wrong grade
wrong scores: [34.5, 44.5, 74.5] · 5 students told they FAILED at 34.5
found by one test at Dipika's desk: 1 failing check, 0 students hurt · found after the cards are posted:
15 reprints, 5 apology letters, a re-marking meeting — the later a bug is found, the more it costs
a passing test proves the code works for THAT input; it cannot prove there is no bug (5/5 passed above)
═══ unit ═══
── the tiny runner: 4 exam questions for grade_v1 (arrange · act · assert)
distinction pass
fail_below_pass_mark pass
just_a_distinction FAIL — expected 'A', got 'B'
out_of_range ERROR — ValueError: score out of range: 101
2 passed, 2 failed · a FAIL is a wrong answer, an ERROR is a crash the test did not expect
── stdlib unittest runs exam/suites/test_grades_unittest.py against grade(): 6 tests · 0 failures · 0 errors
same shape, more help: setUp/tearDown, assertEqual/assertRaises, discovery, -v; pytest is the popular third-party runner
═══ design ═══
── equivalence classes: one question per class is enough to start
-5 → refused 20 → F 40 → D 50 → C 65 → B 90 → A 150 → refused
── boundary values: v1 vs the fixed grade()
34.4: v1 F · fixed F
34.5: v1 F · fixed D ← BUG
35: v1 D · fixed D
44.5: v1 D · fixed C ← BUG
45: v1 C · fixed C
59.5: v1 B · fixed B
60: v1 B · fixed B
74.4: v1 B · fixed B
74.5: v1 B · fixed A ← BUG
75: v1 A · fixed A
every half-mark 0–100 (201 scores): v1 is wrong at [34.5, 44.5, 74.5] — round() sends halves to the EVEN number
── decision table: promoted = average passes AND (attendance ≥ 75% OR a medical note)
average 70 · attendance 90% · note no → promoted
average 70 · attendance 90% · note yes → promoted
average 70 · attendance 60% · note no → held back
average 70 · attendance 60% · note yes → promoted
average 30 · attendance 90% · note no → held back
average 30 · attendance 90% · note yes → held back
average 30 · attendance 60% · note no → held back
average 30 · attendance 60% · note yes → held back
3 conditions → 8 rules → 8 tests; a table shows the rule nobody wrote down
═══ fixtures ═══
── a_student() — a prepared answer sheet with sensible defaults → Katrina · average 75.0 · A · promoted True
.with_attendance(60) → promoted False · + .with_medical_note() → promoted True
.named('Dipika').with_marks(maths=20, science=30, english=40) → average 30.0 · F
── one SHARED answer sheet: order enrol→empty: 1 passed, 1 failed · order empty→enrol: 2 passed, 0 failed
a FRESH sheet per test (setup): 2 passed, 0 failed · reversed: 2 passed, 0 failed · teardown ran 2 times
independent tests pass in any order, alone or together — the order must never be part of the answer
═══ doubles ═══
── the attendance service is far away; stand-in examiners take its place
stub (canned answer) → attendance 86.0% · eligible True
fake (a small working one) → student 42: 86.0% · student 7: 70.0% · eligible False
spy (records calls) → calls ['/attendance/42']
mock (unittest.mock) → a 503 raises AttendanceUnavailable; assert_called_once_with('/attendance/42') holds
── over-mocking: the client now asks '/attendance/42?term=1' — same answers, new URL
fake-based test: eligible True → passes · mock-based test → FAILS (the mock checked HOW, not WHAT)
and a stub can lie: it returned {'days_present': ...} but what if the real service says 'present'? → KeyError 'days_present' in production
double what you do not own or cannot run fast; check the doubles against the real thing (lesson 07)
═══ integration ═══
── Katrina 68 and Dipika 67 in maths; the class average should be 67.5
unit test with a Python fake store → 67.5 ✅ (Python's / keeps the half)
integration test, REAL SQLite in memory, v1 SQL SUM/COUNT → 67 ❌ (INTEGER / INTEGER drops it)
fixed SQL AVG(marks) → 67.5 ✅
save Katrina's maths twice: fake silently overwrites · SQLite → IntegrityError (the PRIMARY KEY the fake never had)
── exam/suites/test_store_sqlite.py: 3 tests · 0 failing · a fresh :memory: database per test
integration tests catch the bugs that live BETWEEN your code and the real thing: SQL, types, constraints
═══ contract ═══
── the consumer (report cards) writes down what it needs: GET /attendance/42 → 200 with days_present:int, days_total:int
provider v1 as today → ✅ honours the contract
provider v2 adds 'late_days' → ✅ honours the contract
provider v3 renames to 'present' → ❌ missing field 'days_present'
provider v4 sends days_total as text → ❌ 'days_total' is str, expected int
the provider runs the consumer's contract in ITS pipeline — v3 and v4 are stopped before they ship
a new field is fine (the consumer ignores what it did not ask for); renames and type changes are not
═══ pyramid ═══
── one end-to-end run through the teaching app: database → attendance → report card → Aishwarya: average 74.33 · B · attendance 86.0% · promoted True
── a cost MODEL (illustrative numbers): unit 0.005 s, integration 0.1 s, e2e 5 s per test; flake chance 0.01% / 0.05% / 1%
pyramid 400 / 80 / 10 tests → 60.0 s per run · a run hits at least one flaky failure 17% of the time
trophy 150 / 300 / 10 tests → 80.8 s per run · a run hits at least one flaky failure 23% of the time
ice-cream cone 50 / 50 / 200 tests → 1005.2 s per run · a run hits at least one flaky failure 87% of the time
both good shapes keep e2e tests FEW; they disagree on whether most tests are unit or integration
═══ property ═══
── property: for every whole n from 0 to 99, grade(n + 0.5) == grade(n + 1) — halves round UP
grade_v1, seed 1: fails on run 25 at n=34 (34.5) · shrinks 34 → smallest failing score 34.5
grade_v1, seed 3: fails on run 9 at n=74 (74.5) · shrinks 74 → 44 → 34 → smallest failing score 34.5
grade (fixed), seed 1: 100 runs, failures: None
grade_v1 property 'a higher score never gets a lower grade', 300 runs → failures: None
grade property 'a higher score never gets a lower grade', 300 runs → failures: None
grade_v1 passes the second property — round() never goes DOWN as the score goes up; a property finds only what it states
the seed is printed so a failure can be replayed; shrinking hands you a simpler failing case than the first one found
═══ mutation ═══
── WEAK suite (7 questions): line coverage of grade() 12/12 = 100% · mutants killed 4/10
s >= 75 → s > 75 SURVIVED
s >= 35 → s > 35 SURVIVED
s >= 60 → s >= 61 SURVIVED
round_half_up(score) → round(score) SURVIVED
return "A" → return "B" killed
return "F" → return "D" killed
score < 0 or → score < 0 and killed
s >= 45 → s <= 45 killed
score > 100 → score > 101 SURVIVED
s >= 75 → s > 74 SURVIVED
── STRONG suite (21 questions): line coverage of grade() 12/12 = 100% · mutants killed 9/10
the survivor of the STRONG suite: 's >= 75' → 's > 74' — s is a whole number, so this mutant is EQUIVALENT (no test can kill it)
coverage says which lines RAN; mutation testing asks whether any test would NOTICE if they were wrong
═══ flaky ═══
── the wobbly clock on the wall: run each test 1000 times with NO code change
time: due today, handed in within the hour passed 967 · failed 33 → reads the real clock (fails after 23:00)
randomness: 3 unseeded scores, one passes passed 964 · failed 36 → unseeded random data
waiting: sleep 200 ms, then check the job passed 924 · failed 76 → a fixed sleep instead of waiting for done
order: shared ROLL list, all 6 orders of 3 tests → 5 orders fail · fresh ROLL per test → 0 fail
── fixed: an injected clock → passed 1000, failed 0 · a seed → the same 3 scores every run · wait for 'done', not for time
a test that both passes and fails on the same code is flaky — quarantine it, find the variable, pin it down
═══ ci ═══
── TDD for a new rule: a student 2 marks short of 35 WITH a medical note is lifted to 35
RED (test first, no code) → 0/5 pass
GREEN (simplest code) → 5/5 pass
REFACTOR (name the rule) → 5/5 pass
── test selection: a change to the attendance client
runs 3 of 6 suites, 28 of 105 tests: unit/test_attendance, contract/test_attendance_pact, e2e/test_report_flow
a change to grade() runs 69 of 105 tests · merging to main still runs all 105
CI order: lint → unit → integration → contract → e2e — fast checks first, stop at the first red
── the whole picture: questions (unit) → the right questions (design) → fresh sheets (fixtures) → stand-ins (doubles) → real database (integration) → promises (contracts) → the pyramid → properties → mutants → a steady clock → CI
✅ done — every paper marked
sys.settrace hook आहे, coverage.py नाही; mutation tester मध्ये हाताने लिहिलेले 10 mutants आहेत, mutmut ने तयार केलेले नाहीत. खरी tools — pytest, Hypothesis, Pact, Playwright, coverage.py, mutmut — प्रत्येक धड्यात येतात. हे कुठे बसते: CI/CD शाळा हे tests pipelines मध्ये चालवते; API शाळा ते तपासत असलेले contracts design करते.संपूर्ण कोर्स एका कॅनव्हासवर. 4K आवृत्तीसाठी क्लिक करा.
एक git branch = एक कल्पना; branch 04 मध्ये धडे 01–04 आहेत. exam/ ची प्रत्येक run सारखीच असते — randomness ला seed दिलेला आहे.
lesson-01-why-testधडा वाचा →आकृती पहा ↗lesson-02-unit-testsधडा वाचा →आकृती पहा ↗lesson-03-test-designधडा वाचा →आकृती पहा ↗lesson-04-fixturesधडा वाचा →आकृती पहा ↗जेव्हा code इतर गोष्टींशी बोलतो: doubles, खरा database, teams मधील contracts, आणि प्रत्येक प्रकारचे किती tests.
lesson-05-test-doublesधडा वाचा →आकृती पहा ↗lesson-06-integrationधडा वाचा →आकृती पहा ↗lesson-07-contract-testsधडा वाचा →आकृती पहा ↗lesson-08-e2e-pyramidधडा वाचा →आकृती पहा ↗tests चीच परीक्षा: generate केलेले inputs, mutants, flakiness — आणि योग्य वेळी योग्य tests चालवणे.
lesson-09-property-basedधडा वाचा →आकृती पहा ↗lesson-10-coverage-mutationधडा वाचा →आकृती पहा ↗lesson-11-flaky-testsधडा वाचा →आकृती पहा ↗lesson-12-ci-tddधडा वाचा →आकृती पहा ↗प्रत्येक धडा एका क्रमांकित बॉक्स-आणि-बाण आकृतीत — स्वतंत्र पानावरही.
बाकावरच सापडलेला bug एका ओळीत सुधारतो; निकालपत्रे पाठवल्यानंतर सापडला तर 15 पुनर्छपाई लागतात. test काय सिद्ध करतो — आणि काय करू शकत नाही.
दीपिकाने गुणांचे grade बनवणारा code लिहिला. कतरिनाने त्याला 5 सोपे प्रश्न विचारले, आणि पाचही बरोबर आले. मग 1200 विद्यार्थ्यांनी 200 गुणांचा पेपर दिला. 34.5% असलेली मुलगी पास व्हायला हवी होती, पण code ने F सांगितले. 15 report cards चुकीचे आले. दीपिकाच्या डेस्कवर अजून एक प्रश्न विचारला असता तर हे एका मिनिटात पकडले गेले असते.
दीपिकाने grade_v1 लिहिले आणि नजरेने तपासले; ते बरोबर दिसले, म्हणून ते थेट report cards छापायला गेले.
Test म्हणजे code साठी exam question, ज्याचे बरोबर उत्तर आधीच माहीत असते. grade_v1 पाच सोपे प्रश्न सोडवते: 5 पैकी 5.
1200 विद्यार्थी, 200 गुणांचा पेपर: round(34.5) = 34, म्हणून 34.5, 44.5 आणि 74.5 वर 15 report cards चुकीचे येतात.
दीपिकाच्या डेस्कवर एक प्रश्न म्हणजे एक failing check; cards पाठवल्यानंतर 15 reprints आणि 5 माफीपत्रे.
पुढच्या धड्यात खरा exam question: arrange, act, assert, आणि FAIL व ERROR तपासणारा छोटा runner.
Arrange, act, assert — परीक्षेचा एक प्रश्न, एका छोट्या runner मध्ये आणि standard library च्या unittest मध्ये.
चांगल्या exam question च्या तीन पायऱ्या असतात. आधी तयारी: गुण 82 आहेत. मग एकदा विचारा: grade काय? मग answer key शी जुळवा: ते A हवे. कतरिनाने असे 4 प्रश्न लिहिले. दोन pass झाले. एकाचे उत्तर चुकले, तो FAIL. एकाने code crash केला, तो ERROR.
तपासणी म्हणजे print() ओळी, ज्या कोणीतरी डोळ्यांनी वाचायच्या; छापलेला grade गुपचूप बदलला तरी कोणाला कळत नसे.
Unit test score ठरवते (arrange), grade_v1(score) एकदा चालवते (act), आणि answer key शी जुळवते (assert).
कतरिनाचे 4 प्रश्न: 2 pass, 74.5 FAIL (अपेक्षित 'A', मिळाले 'B') आणि 101 ERROR: ValueError out of range.
FAIL म्हणजे चुकीचे उत्तर; ERROR म्हणजे test ला अपेक्षित नसलेला crash. unittest 6 tests चालवते: 0 failures, 0 errors.
पुढचा धडा विचारतो कोणते प्रश्न लिहायचे: equivalence classes, boundary values आणि decision table.
Equivalence classes, boundary values आणि decision tables — आणि 34.5 वरचा rounding bug.
कतरिना प्रत्येक गुणाबद्दल विचारू शकत नाही. म्हणून ती गट करते: 45 ते 59 मधल्या सगळ्या गुणांना C, म्हणून एक प्रश्न पुरेसा. सात प्रश्न, सगळे pass. मग ती कडांवर विचारते: 34.4, 34.5 आणि 35. 34.5 वर code F म्हणतो पण नियम D म्हणतो. पकडले! चुका कडांवरच लपतात.
कतरिनाने मनात आलेले गुण निवडले: 90, 65, 50, 40, 10. सगळे pass झाले, आणि rounding bug लपूनच राहिला.
Equivalence classes प्रत्येक band मधून एक गुण घेतात; boundary values प्रत्येक कडेच्या थोडे खाली, नेमके आणि वर विचारतात.
34.5 वर grade_v1 म्हणते F, दुरुस्त grade म्हणते D. सर्व 201 half-marks पैकी v1 फक्त 34.5, 44.5 आणि 74.5 वर चुकते.
Bugs कडांवर राहतात. Decision table 3 हो/नाही अटींचे 8 नियम करते, म्हणून promotion चा कोणताही प्रकार सुटत नाही.
पुढच्या धड्यात answer sheets तयार होतात: योग्य defaults असलेले builders, आणि प्रत्येक test साठी नवा fixture.
तयार उत्तरपत्रिका: builders, setup आणि teardown, आणि कोणत्याही क्रमाने pass होणारे tests.
परीक्षेपूर्वी office प्रत्येक विद्यार्थिनीसाठी answer sheet तयार करते. कतरिनाकडेही एक तयार विद्यार्थिनी आहे: चांगले गुण आणि 90% attendance. कमी attendance तपासायला ती फक्त तेच 60% करते. एकदा दोन प्रश्नांनी एकच sheet वापरली. पहिल्याने त्यावर ऐश्वर्याचे नाव लिहिले, आणि दुसऱ्याला ती कोरी हवी होती. म्हणून आता प्रत्येक प्रश्नाला नवी sheet मिळते.
प्रत्येक test स्वतःची विद्यार्थिनी हाताने बनवत असे, आणि दोन tests एकच roll list वापरत, ज्यावर पहिलीने आधीच लिहिलेले असे.
Fixture म्हणजे तयार answer sheet. a_student() म्हणजे कतरिना: average 75.0, grade A, promoted True, 90% attendance.
.with_attendance(60) ने promoted False होते; .with_medical_note() जोडल्यावर True. प्रत्येक test एकच गोष्ट बदलते.
एकच shared sheet: enrol → empty केल्यास 1 passed, 1 failed. प्रत्येक test ला नवी sheet: दोन्ही क्रमात 2 passed.
पुढच्या धड्यात दूरच्या attendance service ऐवजी पर्यायी परीक्षक येतात: stub, fake, spy आणि mock.
बदली परीक्षक: stub, fake, spy आणि mock — आणि जास्त mocking मुळे चुकीच्या गोष्टीची test कशी होते.
खऱ्या attendance परीक्षिका दुसऱ्या इमारतीत बसतात आणि कधी कधी रजेवर असतात. म्हणून कतरिना पर्यायी परीक्षक वापरते. Stub कडे एकच card असते: 200 पैकी 172 दिवस. Fake कडे विद्यार्थिनींची छोटी चालणारी वही असते. Spy तिला विचारलेला प्रत्येक प्रश्न लिहून ठेवते. Mock ला एकच ठरलेला प्रश्न अपेक्षित असतो आणि दुसरा आला की ती तक्रार करते.
प्रत्येक test खऱ्या attendance service ला बोलवत असे: हळू, कधी बंद, आणि ती रजेवर असली की test fail होत असे.
Test doubles म्हणजे पर्यायी परीक्षक: stub ठरलेले 172/200 देतो, fake विद्यार्थिनींची छोटी चालणारी dict ठेवतो.
Stub: 86.0%, eligible True. Fake: pupil 7 ला 70.0%, eligible नाही. Spy '/attendance/42' नोंदवतो; 503 वर error.
URL मध्ये '?term=1' आले, उत्तरे तीच: fake test pass होते, mock test FAIL होते. त्याने 'काय' नव्हे, 'कसे' तपासले.
पुढच्या धड्यात खोट्या marks book ऐवजी खरा SQLite database येतो, आणि fake मध्ये कधीच न येणारा bug सापडतो.
memory मधील खरा SQLite database fake ला न सापडलेला bug शोधतो: 67, 67.5 नाही.
कतरिनाला maths मध्ये 68 आणि दीपिकाला 67 मिळाले, म्हणून वर्गाची सरासरी 67.5. Python मधल्या खोट्या marks book ने 67.5 सांगितले. खऱ्या database ने 67 सांगितले! Database मध्ये पूर्ण संख्येला पूर्ण संख्येने भागले की अर्धा भाग गळतो. खोट्या वहीत हा bug कधीच येऊ शकला नसता. फक्त खऱ्या database सोबतच्या test ने तो शोधला.
दीपिकाने class average फक्त Python dict वर तपासले, म्हणून production मध्ये चालणाऱ्या SQL ला कोणीच प्रश्न विचारला नाही.
Integration test तुमचा code खऱ्या गोष्टीसोबत चालवते: इथे memory मध्ये तयार केलेला खरा SQLite database.
कतरिना 68, दीपिका 67: fake म्हणते 67.5, SQLite SUM/COUNT म्हणते 67 (INTEGER / INTEGER), आणि AVG(marks) म्हणते 67.5.
काही bugs code आणि database यांच्यामध्ये राहतात. कतरिनाचे maths दोनदा save: fake overwrite करते, SQLite IntegrityError देते.
पुढचा धडा attendance service न चालवता तपासतो: consumer एक contract लिहितो, जो provider ने पाळायलाच हवा.
consumer त्याला काय हवे ते लिहून ठेवतो; provider प्रत्येक release त्याच्याशी तपासतो.
कतरिनाच्या report cards ना attendance office कडून दोन आकडे लागतात: हजर दिवस आणि एकूण दिवस. ती ते एका card वर लिहून office च्या भिंतीवर लावते. Office आपला form बदलण्यापूर्वी तो तिच्या card शी जुळवून पाहते. नवा रकाना जोडणे चालते. days_present चे नाव बदलणे किंवा 200 अक्षरात लिहिणे नवा form बाहेर जाण्यापूर्वीच पकडले जाते.
Attendance office ने आपला form बदलला आणि नंतर सांगितले; पहिल्याच दिवशी report cards बिघडले.
Consumer-driven contract: report cards स्वतःला काय लागते ते लिहून ठेवतात, days_present: int आणि days_total: int.
Provider आपल्या CI मध्ये contract पुन्हा चालवतो: v1 pass, v2 late_days जोडूनही pass, v3 आणि v4 थांबवले जातात.
नवीन field चालते; नाव बदलणे ('present') किंवा type बदलणे (days_total मजकूर म्हणून) चालत नाही, ship आधीच पकडले जाते.
पुढच्या धड्यात शाळेचा पूर्ण दिवस end to end चालतो, आणि suite मध्ये कोणत्या प्रकारच्या किती tests हव्यात ते पाहतो.
End-to-end tests, आणि वेळ व flakiness मध्ये pyramid विरुद्ध trophy विरुद्ध ice-cream cone.
शाळा तपासायचे तीन मार्ग आहेत. डेस्कवरचा unit प्रश्न काही सेकंद घेतो. Marks book ची integration तपासणी थोडा जास्त वेळ घेते. शाळेचा पूर्ण दिवस सगळे एकत्र तपासतो, पण खूप वेळ लागतो आणि उशिरा आलेली bus तो बिघडवू शकते. म्हणून छोटे प्रश्न भरपूर आणि पूर्ण दिवस थोडेच ठेवा. बहुतेक पूर्ण दिवस म्हणजे वितळणारा ice-cream cone.
Team बहुतेक पूर्ण app मधून click करून तपासत असे; प्रत्येक run ला खूप वेळ लागे आणि उशिरा आलेली bus ते वारंवार मोडत असे.
E2e test सगळे भाग एकत्र चालवते: database → attendance → report card, ऐश्वर्याला 74.33, B, promoted True.
Cost model मध्ये pyramid ला एका run ला 60.0 s लागतात, trophy ला 80.8 s, आणि ice-cream cone ला 1005.2 s.
E2e tests जितक्या जास्त तितक्या flaky failures: pyramid चे 17% runs, trophy चे 23%, ice-cream cone चे 87%.
पुढच्या धड्यात प्रश्न machine लिहिते: एक नियम सांगा, seed वरून 100 inputs तयार करा, आणि shrink करा.
एक नियम सांगा, एका seed पासून 100 inputs generate करा, आणि failure ला सोप्या case पर्यंत shrink करा.
कतरिना एकेक प्रश्नाऐवजी एक नियम लिहिते: .5 ने संपणाऱ्या गुणाला पुढच्या पूर्ण संख्येइतकाच grade मिळायला हवा. एक प्रश्न-यंत्र 100 आकडे निवडून नियम तपासते. 9 व्या प्रयत्नात ते 74 निवडते: 74.5 ला B, पण 75 ला A. मग ते सोपी चूक शोधते: 44, मग 34. उत्तर 34.5, तोच जुना bug.
कतरिना प्रत्येक test score हाताने निवडत असे, म्हणून code ला फक्त तिला सुचलेल्या आकड्यांबद्दलच विचारले जाई.
Property-based testing एक नियम सांगते, grade(n + 0.5) == grade(n + 1), आणि seeded machine 100 random n तपासते.
Seed 3 वर grade_v1 run 9 मध्ये 74.5 ला fail होते; shrinking लहान n वापरून 74 → 44 → 34, म्हणजे 34.5.
हाताने एकही score न निवडता धडा 03 चा bug सापडला; दुरुस्त grade 100 runs pass करते, आणि seed failure पुन्हा दाखवतो.
पुढचा धडा tests नाच गुण देतो: 100% coverage असूनही चुका सुटू शकतात, ज्या mutation testing पकडते.
100% line coverage जी 10 पैकी फक्त 4 mutants मारते — coverage काय चालले ते दाखवते, काय तपासले ते नाही.
कतरिनाला आपला exam paper चांगला आहे का ते जाणून घ्यायचे आहे. 7 प्रश्नांचा तिचा weak paper code ची प्रत्येक line चालवतो: 100%. मग ऐश्वर्या code च्या 10 copies बनवते, प्रत्येकीत एक छोटी पेरलेली चूक. Weak paper त्यातल्या फक्त 4 पकडतो. कडांवरचे प्रश्न असलेला strong paper 9 पकडतो. शेवटची तर खरी चूकच नाही.
Team line coverage वर विश्वास ठेवत असे: tests चालताना प्रत्येक line चालली, तर paper पुरेसा चांगला मानला जाई.
Mutation testing grade() च्या एका copy मध्ये एक छोटी चूक पेरते, ज्याला mutant म्हणतात, आणि कोणती test ती ओळखते ते पाहते.
7 प्रश्नांच्या WEAK suite ला 12/12 line coverage आहे, पण ती 10 पैकी फक्त 4 mutants मारते; 21 ची STRONG suite 9/10 मारते.
Coverage कोणत्या lines चालल्या ते दाखवते, काय तपासले ते नाही. शेवटचा survivor, s > 74, equivalent आहे: कोणतीही test त्याला मारू शकत नाही.
पुढचा धडा code न बदलता pass आणि fail होणाऱ्या tests शोधतो: flaky tests चे हलणारे भिंतीवरचे घड्याळ.
डुगडुगणारे घड्याळ: वेळ, क्रम, randomness आणि वाट पाहणे — पुन्हा चालवून सापडतात, त्यांना पक्के करून सुधारतात.
परीक्षा हॉलमधले घड्याळ हलते. बहुतेक दिवस ते बरोबर असते; काही दिवस ते उडी मारते. Flaky test अशीच असते: code बदलला नाही, पण निकाल बदलला. ऐश्वर्या प्रत्येक test 1000 वेळा चालवते. एक 33 वेळा fail होते, फक्त 23:00 नंतर. दुसरी एक list वाटून घेते आणि 6 पैकी 5 क्रमांत fail होते. ती बदलणारी गोष्ट शोधते आणि fixed clock ने ती बांधून ठेवते.
लाल build हिरवा होईपर्यंत पुन्हा चालवला जाई, आणि त्याच code ने वेगळे उत्तर का दिले हे कोणी विचारत नसे.
Flaky test त्याच code वर कधी pass तर कधी fail होते. ऐश्वर्या प्रत्येक test काहीही न बदलता 1000 वेळा चालवते.
Time 33 वेळा fail (23:00 नंतर), randomness 36, 200 ms sleep 76, आणि shared ROLL वर 6 पैकी 5 test orders fail.
प्रत्येक flake मागे एक बदलणारी गोष्ट असते. ती बांधा: fixed clock ने 1000 passed, 0 failed; seed; 'done' ची वाट पाहा.
पुढच्या धड्यात आधी test लिहिली जाते (red, green, refactor), आणि CI फक्त बदलाला लागणाऱ्या tests चालवते.
Red, green, refactor; बदलाला लागू असलेलेच tests चालवा — आणि संपूर्ण चित्र.
एक नवा नियम येतो. कतरिना आधी 5 प्रश्न लिहिते, आणि पाचही fail होतात, कारण अजून code नाही. हे red. दीपिका सर्वात साधा code लिहिते: 5 पैकी 5 pass, green. मग ती तो नीटनेटका करते आणि तरीही 5 पैकी 5 pass. प्रत्येक बदलानंतर CI hall फक्त त्या बदलाला लागणाऱ्या tests चालवतो: 105 पैकी 28. Main मध्ये जाण्यापूर्वी सगळ्या 105 चालतात.
आधी code येत असे आणि tests नंतर, कधी कधी तर नाहीच; प्रत्येक बदलावर एकतर सगळ्या 105 tests चालत किंवा एकही नाही.
TDD: आधी 5 tests लिहा (red, 0/5), मग सर्वात साधा code (green, 5/5), मग तो नीटनेटका करा (refactor, तरीही 5/5).
Test selection: attendance client मधील बदल 6 पैकी 3 suites, 105 पैकी 28 tests चालवतो; grade() मधील बदल 69 चालवतो.
Red सिद्ध करते की test fail होऊ शकते; CI lint → unit → integration → contract → e2e चालवते, पहिल्या red ला थांबते, main वर सगळ्या 105.
पुढे कुठे: CI/CD school या tests खऱ्या pipelines मध्ये चालवते, आणि API school contracts design करते.