Varvara Arzt

Varvara Arzt

/vɐrˈvarə ɑːrtst/ · var-VAH-ra ARTST

PhD researcher in NLP, TU Wien

I work on how language models acquire natural language, how that knowledge is stored, and whether different types of knowledge (e.g. factual and linguistic) are stored in different ways, including in multilingual settings that go beyond high-resource languages.

Mechanistic interpretability, and how far interpretability methods can themselves be trusted · Model diffing · LLM sycophancy · Multilingual NLP · Language modeling · AI safety

Publications

Peer-reviewed

new! 🦜🏝️

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

Varvara Arzt, Allan Hanbury, and Terra Blevins

To appear in Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026), main conference

Word order preferences in LMs are data-driven, not architectural. Comparing decoder-only LMs across 192 artificial languages and typologically diverse natural ones, the same architecture shows a left-branching preference on artificial languages but an emerging right-branching SVO advantage on natural ones as training data scales, and that advantage tracks resource level and data quality rather than word order itself. This implies LLM adoption may gradually narrow word order diversity.

AnnoHID: LLM-Assisted Annotation Framework for Low-Resource Medical Texts

Annisa Maulida Ningtyas, Guntur Budi Herwanto, Yunita Sari, Rifki Afina Putri, Filip Kovacevic, Alaa El-Ebshihy, Varvara Arzt, and Florina Piroi

ACL 2026, System Demonstrations, pages 683–691, San Diego, California, USA

A human-in-the-loop annotation framework for medical text in low-resource languages. On Indonesian clinical data, LLM pre-annotation raised inter-annotator agreement (κ = 0.76 for NER) and human review lifted F1 from 0.39 to 0.45.

Relation Extraction or Pattern Matching? Unravelling the Generalisation Limits of Language Models for Biographical RE

Varvara Arzt, Allan Hanbury, Michael Wiegand, Gábor Recski, and Terra Blevins

IJCNLP–AACL 2025 (main conference), pages 1463–1484, Mumbai, India

RE models pattern-match rather than generalise. Cross-dataset experiments show relation extraction models fail to transfer even within similar domains, and that higher in-dataset performance often signals overfitting to dataset artefacts. Data quality, not lexical similarity, drives robust transfer, and zero-shot baselines occasionally beat every cross-dataset result.

Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

Varvara Arzt and Allan Hanbury

2nd GenBench Workshop on Generalisation (Benchmarking) in NLP @ EMNLP 2024, pages 120–130, Miami, Florida, USA

RE benchmarks and leaderboards obscure real progress. An audit of widely used benchmarks (TACRED, NYT) documenting severe class imbalance, noisy labels and undocumented construction decisions, and showing that aggregate F1 leaderboards hide per-class failure, with recommendations applicable well beyond RE.

TU Wien at SemEval-2024 Task 6: Unifying Model-Agnostic and Model-Aware Techniques for Hallucination Detection

Varvara Arzt, Mohammad Mahdi Azarbeik, Ilya Lasy, Tilman Kerl, and Gábor Recski

18th International Workshop on Semantic Evaluation (SemEval-2024), pages 1183–1196, Mexico City, Mexico

An ensemble unifying model-agnostic and model-aware techniques for hallucination detection, placing 3rd of all systems in the model-aware track (80.6% accuracy).

Preprints

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, et al. (350+ contributors, including Varvara Arzt)

arXiv:2510.24081, 2025

A participatory commonsense benchmark hand-built by over 350 researchers across 100+ languages, associated with the 5th Multilingual Representation Learning (MRL) workshop; I contributed the Albanian subset. State-of-the-art LLMs do well in aggregate but show accuracy gaps of up to 68% on lower-resource languages.

Talks and posters

Analysis of Word Order Biases in Language Models: A Controlled Investigation Using Artificial and Natural Languages

Varvara Arzt

49th Austrian Linguistics Conference (Österreichische Linguistiktagung, ÖLT 2025), University of Klagenfurt, Austria, December 2025

Information Extraction from German Medieval Charters Abstracts

Sandy Aoun, Varvara Arzt, Daniel Luger, and Georg Vogeler

DH 2024, 19th International Conference of the Alliance of Digital Humanities Organizations, Washington, D.C., USA, August 2024

Exploring the Narratives of Rebetiko via Corpus-Based Analysis

Varvara Arzt

Storytelling, DARIAH Annual Event 2022, Athens, Greece, May 2022

Earlier work in linguistics

Published as V. A. Diveeva; name changed to Arzt in June 2019.

Markirovaniye aktantov dvuchmestnich predikatov v albanskom yazike [Actant marking of two-place predicates in Albanian]

Diveeva, V. A.

Book chapter in S. S. Say (ed.), Valentniye klassi dvuchmestnich predikatov v raznostrukturnich yazikach [Valency classes of two-place predicates], pages 72–88. ILS RAS, Saint Petersburg, Russia, 2018

Four conference talks on Albanian argument marking, labile verbs and experiencer constructions

International Conferences for Philology Students, Saint Petersburg State University (2012–2014), and the Institute of Linguistic Studies, Russian Academy of Sciences (2014)

About

I am a PhD researcher at TU Wien, advised by Terra Blevins and Allan Hanbury. Before my PhD I worked on linguistic typology at the Russian Academy of Sciences, including field work in Albanian-speaking villages of Ukraine and southeastern Albania, and completed a research internship at the Austrian Academy of Sciences developing NLP pipelines for medieval German texts.

Education

Teaching

Since 2023 I have taught NLP, Information Retrieval and programming courses with up to 300 students at TU Wien and the University of Klagenfurt.

Languages

Beyond NLP, computer science, mathematics and linguistics, I love learning languages. I speak Russian (native), German, English, Albanian and Modern Greek fluently, some French, and I am currently learning Indonesian. :)

Contact

I'm always happy to collaborate on projects involving multilingual NLP and mechanistic interpretability, or just to chat about NLP and linguistics. Feel free to reach out.