Welcome
About Me
Hi there! I’m Dominik Meier, a researcher focusing on AI safety and security. I study topics around alignment, model internals, and emerging threats in language models. Feel free to reach out!
Contact Information
- Email:
{surname}@gipplab.com - Google Scholar: Profile
Publications
-
Risky Business: Measuring The Faithfulness-Safety Tension
Dominik Meier*, Luca Joshua Francis*, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp
arXiv preprint (2026) · [arXiv] -
BabelSteering: Multilingual Safety Alignment via English Steering Vectors
Dominik Meier*, Emma V. Stein*, Terry Ruas, Jan Philip Wahle, Bela Gipp
arXiv preprint (2026) · [arXiv] -
Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
Kia-Jüng Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp
arXiv preprint (2026) · [arXiv] -
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
Dominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas, Bela Gipp
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025) · [ACL Anthology] · [Explainer Post] -
Towards Human Understanding of Paraphrase Types in Large Language Models
Dominik Meier, Jan Philip Wahle, Terry Ruas, Bela Gipp
Proceedings of the 31st International Conference on Computational Linguistics (COLING 2025) · [ACL Anthology]
* Equal contribution / Shared first authorship