-
Notifications
You must be signed in to change notification settings - Fork 2.9k
Pull requests: huggingface/trl
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Document that padding_free is currently disabled in DPO
#6972
opened Aug 29, 2026 by
behroozazarkhalili
Collaborator
Loading…
Bump https://github.com/huggingface/doc-builder from 0ab9ea03baf111ed8dd83e88233430b663127368 to 1b16dac5e33043af565fdf4c1b5b0fe81d0891c8
dependencies
Pull requests that update a dependency file
pre_commit
Pull requests that update pre_commit code
#6969
opened Aug 29, 2026 by
dependabot
Bot
Loading…
Fix mixed-image Online DPO server batches
#6968
opened Aug 29, 2026 by
DaoyuanLi2816
Contributor
Loading…
5 of 8 tasks
Stop dependabot from opening ruff pre-commit bumps
#6965
opened Aug 28, 2026 by
albertvillanova
Member
Loading…
Use ruff
select instead of extend-select to keep the rule set explicit
#6964
opened Aug 28, 2026 by
albertvillanova
Member
Loading…
Pin layer_types so the tiny Cohere2 model covers both attention types
#6963
opened Aug 28, 2026 by
albertvillanova
Member
Loading…
1 task
Pin layer_types so the tiny Gemma3 and Olmo3 models cover both attention types
#6962
opened Aug 28, 2026 by
albertvillanova
Member
Loading…
Bump the actions group across 1 directory with 5 updates
dependencies
Pull requests that update a dependency file
github_actions
Pull requests that update GitHub Actions code
#6958
opened Aug 28, 2026 by
dependabot
Bot
Loading…
Bump https://github.com/astral-sh/ruff-pre-commit from v0.13.3 to 0.16.4
dependencies
Pull requests that update a dependency file
pre_commit
Pull requests that update pre_commit code
#6957
opened Aug 28, 2026 by
dependabot
Bot
Loading…
Validate reward model tokenizer matches policy in XPOTrainer/NashMDTrainer
#6955
opened Aug 28, 2026 by
amanyagami
Loading…
5 of 8 tasks
Honor reward_processing_classes in XPO and Nash-MD trainers
#6952
opened Aug 27, 2026 by
22elix3r
Loading…
5 of 8 tasks
Compute the loss in prediction_step so Online DPO evaluation runs
#6950
opened Aug 27, 2026 by
behroozazarkhalili
Collaborator
Loading…
ZeroSyncGRPOTrainer: self-contained minimal trainer, one weight copy, never-idle generation
#6949
opened Aug 27, 2026 by
qgallouedec
Member
•
Draft
Remove the experimental Harbor integration (superseded by OpenEnv)
#6948
opened Aug 27, 2026 by
adithya-s-k
Collaborator
Loading…
Add AsyncGRPO Harbor example: any harness, any sandbox, any Harbor dataset, served through OpenEnv
#6947
opened Aug 27, 2026 by
adithya-s-k
Collaborator
Loading…
Remove the unreachable prepare_multimodal_messages_vllm
#6946
opened Aug 27, 2026 by
behroozazarkhalili
Collaborator
Loading…
Warn when reshaped sampling biases the AsyncGRPO importance ratio
#6944
opened Aug 27, 2026 by
behroozazarkhalili
Collaborator
Loading…
Dequantize the bitsandbytes base before the vLLM weight push
#6922
opened Aug 25, 2026 by
behroozazarkhalili
Collaborator
Loading…
fix(vllm): support IPv6 communicator hosts
#6907
opened Aug 25, 2026 by
yikun-c
Loading…
4 of 8 tasks
Support tool-returned images across VLM architectures
#6906
opened Aug 25, 2026 by
DaoyuanLi2816
Contributor
Loading…
3 of 8 tasks
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-07-29.