Running async RL on my DGX Spark cluster
Running an async RL loop across two DGX Spark machines and finding the optimizer offload that dominated weight sync.
Inference and training walkthroughs: measured results, configuration choices, and the limits of each setup.
Running an async RL loop across two DGX Spark machines and finding the optimizer offload that dominated weight sync.
What I learned serving DeepSeek-V4-Flash for RL rollouts on two DGX Spark machines.