TaxBERT

This repository accompanies the paper: Hechtner, F., Schmidt, L., Seebeck, A., & Weiß, M. (2026). How to design and employ specialized large language models for accounting and tax research: The example of TaxBERT. TaxBERT is a domain-adapated RoBERTa model, specifically designed to analyze qualitative corporate tax disclosures.

Update: We added the following features:

SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5146523 Our study develops, evaluates, and applies TaxBERT, a domain-specific BERT model tailored to tax-related language. TaxBERT is particularly well suited to analysing non-legal textual sources, including corporate disclosures. We show that for relatively straightforward classification tasks, it provides only modest incremental benefits over general-purpose large language models, FinBERT, and traditional bag-of-words approaches. However, it delivers significant performance advantages as tasks require a more fine-grained understanding of specialised terminology and contextual distinctions. We further validate TaxBERT in an applied research setting related to Key Audit Matters (KAMs), demonstrating its practical usefulness for empirical tax research. By documenting the development, validation, and practical implementation of TaxBERT, this study provides accounting and tax researchers with a transparent and replicable blueprint for constructing and applying specialised BERT models to domain-specific research questions.

GitHub: https://github.com/TaxBERT/TaxBERT

If the following Guide/Repository is used for academic or scientific purposes, please cite the paper.

Downloads last month
50
Safetensors
Model size
82.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support