How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

Sarkar, Dipankar

Open source

DOI
10.48550/arxiv.2609.30074
Published
2026
Container
Not recorded
Publisher
arXiv
Open access
yes

Credibility signals

limited evidence Score 43/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.48550/arxiv.2609.30074,
  title = {How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure},
  author = {Sarkar, Dipankar},
  year = {2026},
  doi = {10.48550/arxiv.2609.30074},
  url = {https://doi.org/10.48550/arxiv.2609.30074}
}

RIS

TY  - JOUR
TI  - How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
AU  - Sarkar, Dipankar
PY  - 2026
DO  - 10.48550/arxiv.2609.30074
UR  - https://doi.org/10.48550/arxiv.2609.30074
ER  - 

APA

Dipankar, S. (2026). How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure. https://doi.org/10.48550/arxiv.2609.30074

Source records