1 min readfrom Machine Learning

I have a mid-sized GPU cluster and was thinking about giving free compute [D]

I have built an on-prem GPU cluster, 8 nvidia 16GB GPU's and 256GB CPU RAM, 50TB HDD and several TBs of SSDs. I have used it, and currently use it, for ML/AI research. But that research is not constantly running jobs, sometimes I use it heavily and other times it's idle. I was considering just letting people with qualified use cases run jobs on it SLURM style. I don't know if its enough compute to be useful really. Let me know if it's something you'd be interested in using for your research? what would you actually run in ~200 GPU-hours on 8x16GB cards?

I've found it can handle RLVF pretty well, and I have pretrained models up to 500M parameters on it (research size). But obviously it's no stargate cluster

submitted by /u/redwat3r
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#GPU cluster
#NVIDIA
#ML/AI research
#SLURM
#GPU-hours
#16GB GPU
#Pretrained models
#500M parameters
#RLVF
#On-prem
#Compute
#HDD
#SSD
#Machine Learning
#Artificial Intelligence
#GPU
#256GB RAM
#50TB
#Idle resources
#Stargate