
Topic modeling helped discover candidate codes but could not replace human interpretation
2nd International Conference on Quantitative Ethnography (ICQE 2020)
經審核的文章摘要目前以英文提供。
審核摘要

Using Topic Modeling for Code Discovery in Large Scale Text Data, a 2021 conference paper by Zhiqiang Cai, Amanda Siebert-Evenstone, Brendan Eagan, David Williamson Shaffer, examines whether unsupervised topic models can support the discovery of qualitative codes in large textual corpora. The paper compares machine-generated topics with an existing human coding framework and inspects how topic words and human codes align across documents. The review therefore starts from the paper's actual evidence source and purpose rather than from the visual appeal of its final network.
The analytic move is important because ENA represents relations among coded elements, not merely how often each element appears. Rather than treating a topic label as a finished code, the workflow examines topic-word lists as prompts for analysts who must still define constructs, inclusion rules, and contextual meaning. In a defensible workflow, units define whose or what network is accumulated, conversation boundaries and windows define where proximity can become a connection, and the coding scheme defines which aspects of the source material enter the model. Those decisions determine the estimand before normalization, projection, rotation, or plotting begins. A network line is consequently a modeled connection under a documented specification; it is not a direct photograph of thought, collaboration, identity, or learning.
The authors report that topics often mixed several human codes, and words linked to one human code could appear across multiple topics; concise topic-word lists were more useful as discovery aids than direct replacements for codes. This result is most useful as a relational account: it identifies which coded elements were organized together under the study's data and model choices. It should be read alongside unit-level variation, source excerpts or events, and any reported comparison statistics. Visual distance, line thickness, or an attractive subtraction network alone cannot establish practical importance. When a paper combines network output with qualitative return, experimental contrast, trace evidence, or another analytic view, those components strengthen interpretation because they make competing explanations easier to inspect.
The claim boundary is equally central. Alignment with one coded corpus does not establish that a topic model will recover valid constructs in another domain, language, or preprocessing pipeline. ENA cannot on its own repair a weak sample, an unstable codebook, missing contextual evidence, inappropriate dependence assumptions, or a window that crosses contexts that should remain separate. Nor does dimensional reduction preserve every feature of a high-dimensional connection space. The safest conclusion separates three layers: what was observed or collected, what the specified model represents, and what broader explanation the research design can support. Any transfer to a new population, language, activity, platform, or analytic pipeline requires fresh validation rather than visual analogy.
For ENA.HK readers, the paper's durable contribution is that the paper locates automation at a defensible point in qualitative work: generating inspectable candidates while leaving construct judgment and validation with researchers. A reproducible application should save the source-data provenance, segmentation and ordering rules, unit and conversation fields, code definitions, window and weighting choices, normalization and rotation settings, software version, exclusions, and sensitivity checks. It should also retain a route back from every interpreted edge to the qualitative excerpt, observed event, trace record, image element, or document that generated it. That evidence chain keeps the quantitative model and ethnographic meaning in deliberate contact while preventing a descriptive network pattern from being overstated as a causal or universal finding.


