Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so …
Annuncio fornito da Adzuna. Hoomie aggrega questo contenuto da fonti esterne.