LLM efficiency researcher at IRAFM, University of Ostrava.
I work on cheaper inference, recurrent small models, and Czech-language evaluation. Lab notes on X.
Hugging Face · Scholar · ORCID · X · IRAFM
2D early-exit inference — stop in both depth and the input. Extra 1.4–2.3× over optimal layer-wise early exit on vanilla 3–8B models. Paper · data
Recurrent-Gemma-2-2b — Gemma-2 retrofitted into a Huginn/Raven recurrent core. Depth is a runtime knob (num_steps). Also: Recurrent-Llama-3.2-1B.
BenCzechMark — first comprehensive Czech LLM benchmark (50 tasks, 9 categories, duel scoring). Paper (TACL 2025) · blog
Code and models live under irafm-llm and huggingface.co/irafm-llm. I also teach Large Language Models and Introduction to Deep Learning at the University of Ostrava.
| Year | Title | Venue |
|---|---|---|
| 2026 | Two-dimensional early exit optimisation of LLM inference | preprint |
| 2025 | BenCzechMark: a Czech-centric multitask and multimetric benchmark for LLMs | TACL |
| 2024 | Stealing Brains: from English to Czech language model | IJCCI |
| 2024 | Efficient use of large language models for analysis of text corpora | ICPRAM |
Full list on Google Scholar.

