Item analysis
Question statistics
| Question | n | Difficulty p | Discrimination r_pb | Verdict |
|---|---|---|---|---|
| What probability distribution is modelled in supervised learning, according to the conspect? r_pb=-0.43 < 0.15 — does not separate strong from weak | 6 | 0.83 | -0.43 | low discrimination |
| According to Mitchell's definition, what must a computer program do to be considered to learn? | 6 | 0.83 | 0.79 | OK |
| What is the primary computational advantage of using a conjugate prior in Bayesian inference? | 6 | 0.67 | 0.66 | OK |
| What lesson about artificial intelligence does the conspect draw from the ancient Greek myth of Talos? | 6 | 0.33 | 0.22 | OK |
| Why did the Turing test remain a challenging benchmark for AI for many decades while other benchmarks were dismissed after being solved? | 6 | 0.83 | 0.79 | OK |
Rejection rule (from 5 attempts): a question is flagged when its difficulty p falls outside 0.30–0.85 (share answering correctly) or its point-biserial discrimination drops below 0.15 (correlation with the rest of the attempt score). Retiring and regenerating flagged questions runs offline: uv run python -m app.quality … --regenerate.