What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets

Marco Antonio Stranisci, Christian Hardmeier

Open source

DOI
10.1609/aaai.v40i46.41279
Published
2026-03-14
Container
Proceedings of the AAAI Conference on Artificial Intelligence
Publisher
Association for the Advancement of Artificial Intelligence (AAAI)
Open access
unknown

Credibility signals

uncertain Score 64/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1609/aaai.v40i46.41279,
  title = {What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets},
  author = {Marco Antonio Stranisci and Christian Hardmeier},
  year = {2026},
  journal = {Proceedings of the AAAI Conference on Artificial Intelligence},
  doi = {10.1609/aaai.v40i46.41279},
  url = {https://doi.org/10.1609/aaai.v40i46.41279}
}

RIS

TY  - JOUR
TI  - What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
AU  - Marco Antonio Stranisci
AU  - Christian Hardmeier
PY  - 2026
JO  - Proceedings of the AAAI Conference on Artificial Intelligence
DO  - 10.1609/aaai.v40i46.41279
UR  - https://doi.org/10.1609/aaai.v40i46.41279
ER  - 

APA

Stranisci, M. A., & Hardmeier, C. (2026). What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets. Proceedings of the AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v40i46.41279

Source records