Use large language model to enhance reasoning of another large language model through reward updated GRPO.

Yin Y

Open source

DOI
10.1038/s41598-026-39296-8
Published
2026 Feb 11
Container
Scientific reports
Publisher
Not recorded
Open access
yes

Credibility signals

uncertain Score 53/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1038/s41598-026-39296-8,
  title = {Use large language model to enhance reasoning of another large language model through reward updated GRPO.},
  author = {Yin Y},
  year = {2026},
  journal = {Scientific reports},
  doi = {10.1038/s41598-026-39296-8},
  url = {https://doi.org/10.1038/s41598-026-39296-8}
}

RIS

TY  - JOUR
TI  - Use large language model to enhance reasoning of another large language model through reward updated GRPO.
AU  - Yin Y
PY  - 2026
JO  - Scientific reports
DO  - 10.1038/s41598-026-39296-8
UR  - https://doi.org/10.1038/s41598-026-39296-8
ER  - 

APA

Y, Y. (2026). Use large language model to enhance reasoning of another large language model through reward updated GRPO.. Scientific reports. https://doi.org/10.1038/s41598-026-39296-8

Source records