From a draft by Stanford law professor Julian Nyarko and others:
We conducted a blinded evaluation of short-answer tutoring in contracts courses with sixteen U.S. law professors. Participants created 40 representative questions, wrote answers, and judged 2,918 anonymized comparisons between human and LLM responses. Professors rated LLMs far higher than their peers (average win rate = 75.33%), with models performing similarly to the best instructor. LLM responses were also rarely flagged as harmful (3.53%, vs 12.06% for professors). Preferences for LLM answers were consistent across evaluators and reflected shared professional standards….
Sixteen contracts professors from fourteen U.S. law schools—who all use the same casebook to teach the material—authored questions representative of those asked during office hours. From this pool we curated 40 representative questions spanning four instructional categories (Recall: Case or Code, Recall: Doctrine, Hypotheticals, Policy).
Recall questions—whether relating to a case, code or doctrine—tend to be amenable to answers which can be evaluated against a ground truth, and where argumentative strength is of little importance. In contrast, hypotheticals present a short set of facts and ask how the law should be applied. Together with policy questions, which often center on legal or policy design under heterogeneous preferences, providing a strong answer in this category often relies on displaying careful reasoning, weighing competing arguments and other latent, professional standards of quality—even if the relevant doctrine is now settled.