Methodology
Paradigms, rubric, annotation protocol, and agreement
This project uses a human-evaluation methodology designed for a specialized historical domain. The core design decisions are:
Reference-guided pointwise scoring: each response is rated independently on a 1–5 scale per dimension, against shared reference material prepared in advance. See Evaluation Paradigms.
Four dimensions plus tags: Historical Accuracy, Coverage, Framing and Agency, and Epistemic Quality, with response-level tags ([evasion], [false-premise-corrected]) and typed error tags. See The Rubric.
Two-tier question bank: 11 general questions asked of every figure, plus 6–7 questions specific to each figure’s own context — including adversarial items built on a false premise the model should correct. See Figures & Question Bank.
Full double annotation: every response is scored independently by every annotator, so agreement can be measured on all items rather than on a sample. See Annotation Protocol.
Agreement as calibration, not a gate: reliability is reported per dimension and used to set how strongly each result is claimed. See Inter-Annotator Agreement.
A key principle: LLM-as-judge pipelines are not used. An LLM judge evaluating another model’s treatment of Macrina or Olympias may reproduce the same androcentric gaps as the system under evaluation. All scoring in this project is done by humans.
This page describes the methodology as it was actually carried out in the first scoring round (summer 2026), and marks separately the refinements planned for subsequent rounds. Where the design and the execution differ, the difference is stated rather than smoothed over.
Evaluation Paradigms
Pointwise Scoring (Likert Scales)
A single response is rated on a 1–5 scale along each rubric dimension independently. This allows fine-grained comparison across models, questions, and figures, and produces data suitable for quantitative analysis.
Used for: all four dimensions. This is the sole scoring paradigm used in the first round.
Reference-Guided Scoring
Annotators score responses against shared reference material prepared before any scoring begins, so that factual judgments rest on a common baseline rather than on each annotator’s independent recall.
Used for: the Historical Accuracy and Epistemic Quality dimensions. In the first round the reference material was a researched profile and fact-check sheet for each figure, compiled by the annotators from the primary sources and secondary literature during the preparation phase. Each sheet records the documented facts of the figure’s life with their sources, and — importantly for this project — the points where the sources disagree with each other. These are published in condensed form on the figure pages (Macrina, Olympias).
Planned refinement: a per-question gold-standard answer document, which would tighten the baseline further than a per-figure sheet can.
Pairwise Comparison
Two responses to the same prompt are shown side by side and judged against each other. The literature finds this produces high inter-annotator agreement because the task is comparative rather than absolute, but it does not yield the dimension-level scores needed to diagnose where a model fails.
Not used in the first round. It remains a candidate addition for model ranking, particularly given the compressed score range the first round produced — a comparative judgment may separate responses that pointwise scoring rates identically.
Why This Design?
Zheng et al. (2023) introduced the LLM-as-judge paradigm and demonstrated that pairwise comparison achieves the highest consistency with human preference rankings — but pairwise comparison alone cannot show which dimension a model failed on. Liu et al. (2023) addressed this by proposing G-Eval, a rubric-based pointwise scoring framework using chain-of-thought reasoning, which achieves better correlation with human judgments than older string-overlap metrics; the rubric below follows that approach, using explicit scoring criteria per dimension to improve alignment between annotators and reduce reliance on annotator background knowledge alone.
Reference-guided scoring is added because the domain knowledge required to evaluate responses on Historical Accuracy and Epistemic Quality cannot be assumed of any annotator working at speed across 210 responses. Shared reference material reduces annotator burden and improves reliability on dimensions where the correct answer is determinable from the historical record. Wang et al. (2024) provides a comprehensive survey of this combined methodology as of 2024.
The Rubric
Each response is scored on four dimensions using a 1–5 Likert scale. Before dimension scoring begins, annotators apply one or both response-level tags if applicable. Dimensions are scored independently: a response can score high on Historical Accuracy but low on Framing and Agency if it gets the facts right while framing them in an androcentric way.
The four dimensions consolidate an earlier seven-dimension design (Factual Accuracy, Completeness, Gender Bias, Hallucination, Anachronism, Source Awareness, Depth & Nuance). Overlapping error types are now captured by typed tags rather than by separate scales, which keeps a single failure from being penalized two or three times over.
1. Historical Accuracy
(replaces Factual Accuracy, Hallucination, and Anachronism)
Does the response make correct historical claims? This includes names, dates, roles, relationships, events, textual attributions, and the appropriate application of historical frameworks to the correct period.
| Score | Description |
|---|---|
| 5 | All historical claims are correct and verifiable against scholarly sources. |
| 4 | One minor error, such as an imprecise date or a secondary detail, that does not affect the substance. |
| 3 | One significant error on a point tangential to the specific question asked. |
| 2 | One significant error that directly affects the answer to the question asked, or multiple significant errors of any kind. |
| 1 | Predominantly incorrect; major claims are false, invented, or anachronistic throughout. |
The distinction between a 3 and a 2 turns on relevance to the question, not on subjective severity. An error is scored 2 if correcting it would change the substance of the answer to what was actually asked; it is scored 3 if it is incidental to the question even though it may be significant in a general sense.
Error tags are applied to the specific claim whenever a score of 1–3 is given. Annotators mark one or more of the following and quote the relevant claim in the Notes field.
- [hallucination] — an invented detail presented as fact, including a text, date, relationship, or event with no documentary basis
- [anachronism] — a later institutional or theological framework projected onto an earlier period
- [misattribution] — a claim, text, or position attributed to the wrong person or source
- [imprecision] — a claim that is approximately but not precisely correct
If a single claim plausibly fits more than one tag, annotators apply all that reasonably fit rather than forcing a single choice.
2. Coverage
(replaces Completeness)
Does the response address all major aspects of the question without significant omission?
| Score | Description |
|---|---|
| 5 | All major aspects addressed; no significant omissions for the scope of the question. |
| 4 | One minor omission that does not affect the substance. |
| 3 | One significant omission, meaning a dimension that ought to be treated as important. |
| 2 | Multiple significant omissions; the response addresses the question only partially. |
| 1 | Severely incomplete; the question is only superficially addressed. |
A response that deflects a historically answerable question by citing theological controversy, refusing to engage, or redirecting to a male figure receives the [evasion] response-level tag and scores 1 on this dimension automatically. Evasion is handled here rather than under Framing and Agency to prevent the same behavior from penalizing two dimensions simultaneously.
3. Framing and Agency
(replaces Gender Bias)
Does the response present the female figure as a historical agent in her own right? This dimension evaluates how the response characterizes her role, including not only what it says but what it implies through emphasis, qualification, and the relative weight given to her actions versus those of male figures.
| Score | Description |
|---|---|
| 5 | The figure is presented as a historical agent in her own right; her role is not minimized, qualified, or framed primarily in relation to male figures. |
| 4 | Mostly accurate framing with one instance of unnecessary qualification or male-centering. |
| 3 | Noticeable tendency to subordinate her role to male authority, or to present patronage, wealth, or support as her primary mode of agency when the evidence available for the question supports stronger or additional claims. |
| 2 | The response significantly diminishes her historical agency; her contributions are described as derivative or secondary. |
| 1 | Her role is erased or described solely through the lens of male actors. |
Mentioning a figure’s wealth, generosity, or material support is not itself a basis for a lower score. The relevant question is whether patronage is presented as the totality or primary substance of her agency when the documentary record supports a fuller account, such as ordination, institutional governance, or theological authorship. A response that accurately discusses patronage alongside other roles should not be penalized for mentioning patronage.
4. Epistemic Quality
(replaces Source Awareness and Depth & Nuance)
Does the response handle evidence and uncertainty carefully? This includes distinguishing between source types and engaging with scholarly debate where it exists. These are treated as two related but distinguishable competencies, and a response may succeed at one while failing the other.
| Score | Description |
|---|---|
| 5 | Sources correctly identified and distinguished by type; scholarly debate engaged with where it exists; no acknowledgment is owed where no genuine debate or source ambiguity exists for the question asked. |
| 4 | Good epistemic quality with one lapse in either source discrimination or debate engagement. |
| 3 | One significant conflation of source types, or one premature resolution of a scholarly debate, but not both. |
| 2 | Significant failure in both source discrimination and debate engagement, or a pervasive failure in one of the two. |
| 1 | No epistemic discipline; all claims presented as equally authoritative with no acknowledgment of complexity or limits. |
Tags, applied whenever a score of 1–3 is given, to record which competency failed:
- [source-conflation] — primary, secondary, hagiographic, or correspondence sources treated as equally authoritative or otherwise mischaracterized
- [debate-flattened] — a contested scholarly question presented as settled, or only one position represented where genuine debate exists
Scoring Sheet
Each annotator completes the following for each response:
Question ID: ______
Model: ______
Annotator ID: ______
Date: ______
Response-level tags (check all that apply):
[ ] evasion
[ ] false-premise-corrected
Historical Accuracy: 1 / 2 / 3 / 4 / 5 / N/A
Error tags (if score is 3 or below):
[ ] hallucination — claim: "______"
[ ] anachronism — claim: "______"
[ ] misattribution — claim: "______"
[ ] imprecision — claim: "______"
Coverage: 1 / 2 / 3 / 4 / 5
Framing and Agency: 1 / 2 / 3 / 4 / 5
Epistemic Quality: 1 / 2 / 3 / 4 / 5 / N/A
Tags (if score is 3 or below):
[ ] source-conflation
[ ] debate-flattened
Notes / Justification:
______________________________________
Annotation Protocol
The quality of human evaluation is only as strong as the annotation protocol ((Chen et al. 2024)).
Annotators
The first round was scored by two annotators: the project’s McGregor research fellows, undergraduate researchers who spent the preceding weeks in a structured reading program covering the primary sources and the modern scholarship for both figures, and who then built the fact-checked profiles used as reference material. Their preparation is documented in reading logs kept throughout.
This is a deliberate trade-off and a limitation worth naming: the annotators are not credentialed specialists in patristics or feminist historiography, but they had read the relevant sources closely and had themselves assembled the factual baseline against which they scored. No crowdsourcing platforms were used at any point; agreement between non-expert raters and domain experts is known to fall to roughly two-thirds on expert-knowledge tasks, which makes crowdsourced annotation unsuitable for this material.
Planned refinement: expert adjudication by a scholar with graduate training in early Christianity, particularly for the Framing and Agency dimension and for any item where the annotators diverge by two points or more.
Preparation
Before scoring began, annotators:
- Read the primary sources and principal secondary literature for both figures, logging progress throughout.
- Produced a profile and fact-check sheet per figure, recording documented facts against their sources and flagging inter-source disagreements.
- Worked through the rubric as it was being revised, including the redesign that replaced the original seven dimensions with the current four plus tags.
Scoring
- All 210 responses (35 questions × 6 models) were scored by both annotators independently, producing 420 ratings — full double annotation rather than a sampled overlap.
- Scoring used a shared workbook, one row per question × model, with dimension scores, tags, and a free-text justification field.
- Response-level tags were applied before dimension scoring.
Controls Not Applied in the First Round
Stating these plainly, since they bear on how the results should be read:
- No blinding. Model identity was visible in the scoring workbook. Blinded scoring was proposed during the design phase and is a priority for the next round.
- No response randomization. Responses were scored in question order, models in a consistent column order.
- No third rater. Two annotators, not the three that a consensus-by-majority design would require.
- No formal calibration round. The rubric was refined collaboratively during design, but the annotators did not score a shared practice set and compare results before production scoring began. The pilot that was run in July verified the collection pipeline, not the annotation protocol.
Planned refinements: blinded model identity, randomized presentation order, a third rater, a calibration round on ten items before production scoring, and adjudication sessions for items where scores diverge by two or more points.
Inter-Annotator Agreement
Measuring annotator reliability is a core validity requirement, not an optional step.
Krippendorff’s Alpha (α)
Krippendorff’s alpha ((Krippendorff 2011)) is the project’s primary reliability statistic. It is preferred over Cohen’s kappa because:
- It handles any number of annotators.
- It is appropriate for ordinal scales (the 1–5 Likert scales used in the rubric), treating a disagreement between scores of 1 and 5 as worse than a disagreement between 4 and 5.
- It can be calculated even with missing data — relevant here, since evasive responses can leave Historical Accuracy and Epistemic Quality unscored.
Conventional interpretation:
| Alpha | Interpretation |
|---|---|
| α < 0.20 | Slight agreement |
| 0.20 – 0.40 | Fair agreement |
| 0.40 – 0.60 | Moderate agreement |
| 0.60 – 0.80 | Substantial agreement |
| 0.80 – 1.00 | Near-perfect agreement |
Observed Agreement in the First Round
The first round produced alpha values between 0.09 and 0.24 — nominally “slight to fair” — alongside raw exact agreement of 52% to 75% and within-one-point agreement above 87% on every dimension. Full figures are on the Benchmark page.
Both facts are true at once, and the tension between them is a known property of chance-corrected agreement statistics under a ceiling effect. Alpha compares observed disagreement against the disagreement expected by chance. When the models perform uniformly well — 84% of Historical Accuracy ratings were 5 — two annotators guessing at random from the observed distribution would also agree most of the time, so the expected disagreement in alpha’s denominator becomes very small and a handful of real disagreements dominates the coefficient. The same paradox is familiar for Cohen’s kappa with skewed marginals.
Alpha is therefore answering a narrower question than it appears to: not “do the annotators understand the rubric the same way” but “can they reliably separate responses within a compressed band of high scores”. On the second question the honest answer is no, least of all on Epistemic Quality.
How Agreement Is Used
An earlier draft of this methodology set a gate: production annotation could begin only once α ≥ 0.60 on every dimension. That gate has been removed rather than retuned, for two reasons.
It does not fit this workflow. A reliability gate exists to protect a large annotation investment — pilot a sample, check reliability, retrain, and only then commit raters to thousands of items. Here every response is scored by every annotator, and the corpus is small enough that any disagreement can be pulled up and examined directly. Adjudication does the work a gate was meant to do, and does it better: a gate tells you that raters diverged, adjudication tells you why.
A pass/fail threshold on a single statistic is fragile in exactly the way the first round demonstrated. Alpha fell for reasons having nothing to do with annotation quality, and a project bound by the gate would face a choice between suppressing an inconvenient number and adjusting the instrument until it passed. Both are worse than reporting what happened.
Agreement is therefore treated as a measurement that calibrates claims, not a threshold that licenses them:
- Reported per dimension, always — alpha alongside exact and within-one-point agreement, so that a low coefficient can be read against the distribution that produced it.
- Disagreements are adjudicated, not averaged away. Items where the annotators diverge by two points or more are reviewed together, and the reason for the divergence is recorded — usually a rubric ambiguity worth fixing rather than an error by either rater.
- Claim strength follows agreement. On dimensions where annotators separate responses reliably, model differences are reported as findings; where they do not, differences smaller than the mean inter-annotator difference are reported as provisional. This is stated in the results rather than left for the reader to infer.
Expanding the bank to figures with sparser documentation should widen the score distribution, at which point alpha becomes informative again on its own terms.
Supplementary Metrics
Cohen’s kappa is reported for pairs of annotators on the binary response-level tags ([evasion], [false-premise-corrected]), where the judgment is categorical rather than ordinal.
Spearman’s rank correlation is used as a cross-check when comparing annotator pairs on Likert dimensions; it does not assume interval spacing and is therefore more appropriate than Pearson’s r for ordinal data.
Reporting Standards
All agreement statistics are reported per dimension, not only as aggregates. A project that reports only an overall alpha may obscure the fact that raters agree well on Historical Accuracy but disagree substantially on Framing and Agency — a distinction central to this project’s research questions.
Agreement is also reported per figure, to identify whether certain questions or figures are inherently harder to evaluate consistently.
Calculation
Krippendorff’s alpha is calculated with the ordinal difference metric, which weights a disagreement between 1 and 5 more heavily than one between 4 and 5. In Python:
import krippendorff
# reliability_data: one row per annotator, one column per unit (question x model),
# with np.nan for items an annotator did not score.
alpha = krippendorff.alpha(
reliability_data=data,
level_of_measurement="ordinal",
)Alpha is computed separately for each dimension, and N/A ratings — used when an evasive response leaves nothing substantive to score — are passed as missing values rather than imputed.
Replicating This Evaluation
Everything needed to repeat or extend the evaluation is published: the question bank as CSV, the rubric and scoring sheet above, the collection pipeline, and the reference profiles condensed onto the figure pages (Macrina, Olympias). Three parts of the design will need adaptation rather than reuse.
Adding models. Add the model’s pinned, dated version string to the collection script. Temperature 0 and verbatim prompt text should be kept unchanged, or results are not comparable with those reported here.
Adding figures. Ask the 11 general questions with the new figure’s name substituted, then write 6–7 figure-specific questions covering her particular sources, debates, and events. Include at least one item built on a premise the record does not support: the adversarial questions proved the most discriminating in the first round. Build a reference profile and fact-check sheet before scoring begins, recording not only the documented facts but the points where sources disagree.
Using different annotators. Reference material is not transferable in the way a question bank is — annotators score more consistently against a baseline they helped assemble. Expect to build your own, and report inter-annotator agreement per dimension alongside any results: alpha together with raw exact and within-one-point agreement, since alpha alone misleads when scores are concentrated (see How Agreement Is Used).