Status: v0.1.0, in active development. Feedback welcome.
craft operationalizes the C-R-A-F-T framework of Ko, Tai, and Webb Williams into 12 short R functions. It does not call LLM APIs itself: it takes the outputs you already have — labels, confidences, rationales — and walks them through five steps.
Step
Functions
What it does
Construct
role()
Documents how the LLM is being used (annotator, ML system, silicon participant) and which metric family applies.
If you use craft, please cite both the framework paper and the package:
Ko, Hyein, and Yuehong Cassandra Tai. “Can We Trust LLM-Generated Data? The CRAFT Framework for Measurement and Inference in Political Science.” Under review.
Tai, Yuehong Cassandra, and Hyein Ko. 2026. craft: A CRAFT Pipeline for Evaluating LLM-Generated Data. R package version 0.1.0. https://github.com/casstai/craft-r
Large language models (LLMs) are increasingly used for text annotation in social science, but standard performance metrics do not explain why errors occur. Low F1 may reflect limitations of the annotation system, defects in the written codebook, or inconsistency in reference labels. We propose a pre-deployment codebook audit for LLM-assisted annotation. The audit uses class-specific metrics and confusion matrices to identify problematic classes and boundaries, structured document-level review to attribute error sources, and targeted codebook revision followed by held-out evaluation.
This project asks what large language models actually learn when they are given a codebook, a set of labeled examples, or fine-tuning data for an annotation task — and whether the resulting labels reflect the construct a researcher intended.
Large language models are increasingly used for data generation in political and social science, yet the discipline lacks a shared standard for validating their output. Existing frameworks address pieces of the workflow, mostly covering a single stage. We propose C-R-A-F-T, a five-step framework that connects construct definition through inferential adjustment within a single, model-agnostic specification: C-onstruct roles and tasks; R-eport dual-track metrics; A-ssess stability across prompts, and models; F-ield human audit and adjudication; T-ranslate to inference incorporating uncertainty.
We introduce the Digitally Accountable Public Representation (DAPR) Database, an innovative archive that systematically tracks and analyzes the online communications of federal, state, and local elected officials in the U.S. Focusing on X/Twitter and Facebook, the current database includes 28,834 public officials, their demographic information, and 5,769,904 Tweets along with 450,972 Facebook posts, dating from January 2020 to December 2024. The database integrates three interconnected datasets: metadata on elected officials, weekly aggregated X data, and weekly aggregated Facebook data.
Elected officials increasingly communicate across multiple social media platforms. We argue that platforms are not neutral channels: the same political pressures can be expressed differently across them. Using 6.19 million Facebook and X posts from 6,356 U.S. state legislators, we examine climate attention and stance. Ideology structures climate communication differently across platforms: in 2020–2021, liberal legislators devote substantially more climate attention on X, whereas ideological differences in climate stance are stronger on Facebook.
Prominent recent works have measured democratic support using a single latent variable that purports to span a single dimension from steadfast opposition to whole-hearted support. This ignores ample evidence that support for democracy is complex and multidimensional. Here we provide a series of validation tests of the sort of cross-national time-series latent variable measures employed in recent research by reference to questions on support for liberal democracy and opposition to its erosion from multiwave surveys conducted around the world.
Prominent recent work argues that support for democracy behaves thermostatically—that democratic erosion boosts democratic support while deepening democracy yields public backlash—and further contends that there is no evidence for the classic argument that democracy itself increases democratic support over time. Here, we document how these conclusions depend on subtle choices in measurement coding that constitute “researcher degrees of freedom”: analyses employing alternative reasonable choices provide little or no support for the original conclusions.
How does public climate change concern and policy preferences affect climate policy outputs? Prevailing explanations regarding the adoption of climate change (mitigation) policies often focus on collective action, institutions, distributive politics, and policy diffusion. Despite being the focus of many studies on public policy in other issue areas, the role of public opinion on national climate policy has not been studied simply because of a lack of survey questions repeated consistently across years and countries.
Do democratic regimes depend on public support to avoid backsliding? Does public support, in turn, respond thermostatically to changes in democracy? Two prominent recent studies (Claassen 2020a; 2020b) reinvigorated the classic hypothesis on the positive relationship between public support for democracy and regime survival—and challenged its reciprocal counterpart—by using a latent variable approach to measure mass democratic support from cross-national survey data. However, both studies used only the point estimates of democratic support.