I notice that the memory usage for RevGAN seems very high for a network utilizing reversible networks.
With 11 GB of memory I was only able to have a batchSize of 10 (15 was too high).
I checked that use_naive = False, so it should be using the less memory intensive method of doing backpropagation through the generator.
Is it because the discriminators are eating all the memory at this point or what is going on here? I was hoping to be able to have a batch size of around 500.
I'm currently testing it on satellite-maps images with the following settings:
--dataroot ./datasets/maps --name example_run --model unpaired_revgan --batchSize 15
Which gives me:
`----------------- Options ---------------
D_rollout: 1
batchSize: 15 [default: 1]
beta1: 0.5
checkpoints_dir: ./checkpoints
continue_train: False
coupling: additive
dataroot: ./datasets/maps [default: None]
dataset_mode: unaligned
deconv: transposed
display_freq: 50
display_id: -1
display_ncols: 4
display_port: 8097
display_server: http://localhost
display_winsize: 256
epoch_count: 1
fineSize: 256
gpu_ids: 0
grad_reg: 0.0
init_gain: 0.02
init_type: normal
input_nc: 3
isTrain: True [default: None]
lambda_A: 10.0
lambda_B: 10.0
lambda_identity: 0.5
loadSize: 286
lr: 0.0002
lr_decay_iters: 50
lr_policy: lambda
max_dataset_size: inf
model: unpaired_revgan [default: cycle_gan]
mute_optprint: False
nThreads: 4
n_downsampling: 2
n_layers_D: 3
name: example_run [default: experiment_name]
ndf: 64
ngf: 64
niter: 100
niter_decay: 100
no_flip: False
no_html: False
no_lsgan: False
no_output_tanh: False
norm: instance
output_nc: 3
phase: train
pool_size: 50
print_freq: 50
resize_or_crop: resize_and_crop
save_epoch_freq: 25
save_latest_freq: 5000
serial_batches: False
suffix:
update_html_freq: 50
use_naive: False
verbose: False
which_direction: AtoB
which_epoch: latest
which_model_netD: basic
which_model_netG: resnet_9blocks
----------------- End -------------------
len(A),len(B)= 1096 1096
dataset [UnalignedDataset] was created
#training images = 1096
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
model [UnpairedRevGANModel] was created
---------- Networks initialized -------------
[Network G_A_enc] Total number of parameters : 0.378 M
[Network G_core] Total number of parameters : 5.313 M
[Network G_A_dec] Total number of parameters : 0.378 M
[Network G_B_enc] Total number of parameters : 0.378 M
[Network G_B_dec] Total number of parameters : 0.378 M
[Network D_A] Total number of parameters : 2.765 M
[Network D_B] Total number of parameters : 2.765 M
create web directory ./checkpoints\example_run\web...`
I notice that the memory usage for RevGAN seems very high for a network utilizing reversible networks.
With 11 GB of memory I was only able to have a batchSize of 10 (15 was too high).
I checked that use_naive = False, so it should be using the less memory intensive method of doing backpropagation through the generator.
Is it because the discriminators are eating all the memory at this point or what is going on here? I was hoping to be able to have a batch size of around 500.
I'm currently testing it on satellite-maps images with the following settings:
--dataroot ./datasets/maps --name example_run --model unpaired_revgan --batchSize 15
Which gives me:
`----------------- Options ---------------
D_rollout: 1
batchSize: 15 [default: 1]
beta1: 0.5
checkpoints_dir: ./checkpoints
continue_train: False
coupling: additive
dataroot: ./datasets/maps [default: None]
dataset_mode: unaligned
deconv: transposed
display_freq: 50
display_id: -1
display_ncols: 4
display_port: 8097
display_server: http://localhost
display_winsize: 256
epoch_count: 1
fineSize: 256
gpu_ids: 0
grad_reg: 0.0
init_gain: 0.02
init_type: normal
input_nc: 3
isTrain: True [default: None]
lambda_A: 10.0
lambda_B: 10.0
lambda_identity: 0.5
loadSize: 286
lr: 0.0002
lr_decay_iters: 50
lr_policy: lambda
max_dataset_size: inf
model: unpaired_revgan [default: cycle_gan]
mute_optprint: False
nThreads: 4
n_downsampling: 2
n_layers_D: 3
name: example_run [default: experiment_name]
ndf: 64
ngf: 64
niter: 100
niter_decay: 100
no_flip: False
no_html: False
no_lsgan: False
no_output_tanh: False
norm: instance
output_nc: 3
phase: train
pool_size: 50
print_freq: 50
resize_or_crop: resize_and_crop
save_epoch_freq: 25
save_latest_freq: 5000
serial_batches: False
suffix:
update_html_freq: 50
use_naive: False
verbose: False
which_direction: AtoB
which_epoch: latest
which_model_netD: basic
which_model_netG: resnet_9blocks
----------------- End -------------------
len(A),len(B)= 1096 1096
dataset [UnalignedDataset] was created
#training images = 1096
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
[0]
initialize network with normal
model [UnpairedRevGANModel] was created
---------- Networks initialized -------------
[Network G_A_enc] Total number of parameters : 0.378 M
[Network G_core] Total number of parameters : 5.313 M
[Network G_A_dec] Total number of parameters : 0.378 M
[Network G_B_enc] Total number of parameters : 0.378 M
[Network G_B_dec] Total number of parameters : 0.378 M
[Network D_A] Total number of parameters : 2.765 M
[Network D_B] Total number of parameters : 2.765 M
create web directory ./checkpoints\example_run\web...`