Methods & limitations
Show the work.
Keep the caveats.
What the archive contains, what the labels mean, and where this analysis stops.
1. The collected record
This release contains 16,526 posts from @jk_rowling, dated 17 September 2009 through 15 August 2026. It is a snapshot of a research archive frozen in August 2026. It includes original posts, replies, reposts, and 93 records marked as recovered deleted posts by the source archive.
The archive does not establish that every post published by the account has been collected. A missing post, an empty period, or an unavailable context does not prove that no post existed. Deleted status is inherited from the archive and has not been checked against the live platform for this release.
2. What the classifications mean
16,471 posts have classifications; 55 do not. This website reuses the existing research annotations. It does not run a new classifier or infer labels for the remaining posts.
The stored model identifier is claude-sonnet-5. The original pipeline supplied post text and available quoted-post or reply metadata, requesting structured topic, stance, rhetorical-device, descriptor, and intensity labels. Available context varied. Context recovered later and displayed here was not necessarily available to the original classifier.
Trans-related means the post concerns trans people, transition, sex or gender definitions, or associated debates about law, sport, healthcare, spaces, and safeguarding. It is a topic label. It does not mean that every included post is hostile, and it does not assert the identity of anyone discussed in a post.
| Label | Definition used by the coding scheme |
|---|---|
| Hostile | Derogatory or dehumanizing. |
| Mocking | Ridicule or contempt. |
| Critical | Opposition without derogatory language. |
| Defensive | Defending the author or allies from criticism. |
| Neutral | No supportive or opposing stance assigned. |
| Supportive | Support toward trans people or their rights. |
Intensity is the model’s ordinal 0–4 rating of hostility or derogation; it is not an objective measurement. Confidence is the model’s self-reported score, not a calibrated probability of accuracy. Topic, descriptor, and rhetorical-device labels are multi-label annotations.
“Misgendering / sex denial” is the reader-facing name for the source category sex_denial. Other source descriptor categories are sexualization, political framing, jargon/acronym, delegitimizing language, derogatory slurs, and neutral wording. These are interpretations of language in context, not independently adjudicated facts.
3. Review and corrections
The export uses v_analysis, the corrected view of the classifications, which gives stored overrides precedence over the original labels.
Two overrides are attributed to the human reviewer. The other 63 came from the stored Opus early-years recheck: 23 were marked “no” and 40 “unclear.” The 40 unclear cases were subsequently excluded through the review process. An unclear finding is not a demonstrated false positive.
These corrections remove unsupported trans-related classifications from the aggregate counts. The archive’s “Corrected labels” filter exposes the affected records. Removed topic labels are not presented as positive evidence elsewhere on this site.
The corrected dataset retains one supportive classification. You can read the post and its record. Earlier exports and narrative copy may have reported zero; this release follows the corrected source data.
The historical dashboard reported agreement with a stronger model on a sample. Model-to-model agreement is not independent human validation, so this site does not present that figure as an accuracy guarantee. No claim is made that every classification has been reviewed by a person.
4. Counting rules
- Annual topic share: trans-related classified posts divided by all classified posts in the same UTC calendar year. Unclassified posts are excluded from both numerator and denominator. Replies and reposts are included.
- Stance distribution: one stance per trans-related classified post. “Critical, mocking, or hostile” sums those three categories; it does not collapse their definitions.
- Topics, devices, and descriptors: a post counts once in each assigned category. Multiple categories can apply, so percentages need not sum to 100%.
- Engagement: the existing quarterly study counts original posts, excluding replies and reposts. It uses median likes recorded at collection. It does not measure unique people, impressions over time, or causal effects.
- Partial periods: 2026 ends in August. It is not a full year, and its count should not be compared directly with complete years. Early years have very small samples.
5. What this cannot establish
The archive is subject to collection gaps and survivorship bias. The recorded reply or quote context may be missing, incomplete, or recovered after classification. Sarcasm and brief replies are especially difficult to classify. A label alone cannot substantiate a factual allegation about a person mentioned in a post.
Engagement counts are snapshots. Follower growth, platform changes, post age, and deletion can all affect comparisons. An increase in likes does not establish that a particular stance caused that increase.
Language quoted from a post is Rowling’s or the original quoted author’s language. Its inclusion is not an endorsement. Original links may no longer work; a Wayback lookup is a search for a capture, not a guarantee that one exists.
6. Editorial position
“J.K. Rowling is transphobic” is this site’s editorial assessment of her public rhetoric. The numerical research measures topics and coded language; it does not purport to measure private beliefs or motivations. Readers can inspect the underlying material and disagree with the interpretation.
The homepage examples are selected reference points that show different labels, including supportive language. They are not a representative statistical sample. The archive includes all collected posts for this account rather than only the examples used on the homepage.
To make a correction request precise, use the “Copy record link” control on the post page and identify the text, classification, or missing context at issue. Each record includes the source post ID and a link to the original.
7. Data and provenance
This release was rebuilt from the existing local corpus on 19 September 2026. Public exports include only this account’s posts and available public context. Raw API payloads, internal reviewer notes, and other account timelines are not part of the release.
Source lineage: the original Rowling classification pipeline and dashboard; the corrected v_analysis view in the corpus; and the existing engagement_jkrowling.csv engagement study. Tweet IDs remain strings to preserve their full precision.