I am using the default config, and also running on a system with 8 A100 GPUs (each 40GB vRAM). But i still get the following error:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 80.00 MiB. GPU 1 has a total capacity of 39.50 GiB of which 17.38 MiB is free. Including non-PyTorch memory, this process has 39.47 GiB memory in use. Of the allocated memory 39.06 GiB is allocated by PyTorch, and 8.91 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
How much memory is needed to run this? and is there anyway to distribute things better across multiple GPUs? are there any more lightweight versions of the base models i could use to get this running?
I am using the default config, and also running on a system with 8 A100 GPUs (each 40GB vRAM). But i still get the following error:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 80.00 MiB. GPU 1 has a total capacity of 39.50 GiB of which 17.38 MiB is free. Including non-PyTorch memory, this process has 39.47 GiB memory in use. Of the allocated memory 39.06 GiB is allocated by PyTorch, and 8.91 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)How much memory is needed to run this? and is there anyway to distribute things better across multiple GPUs? are there any more lightweight versions of the base models i could use to get this running?