Approx ~131 million total. Made up of estimated: 102M - modified BERT-base-Chinese text encoder 26M - 3D U-Net-style vision/anatomy encoder 2.8M - projection layers, anatomy-specific projections, query tokens and…
Approx ~131 million total. Made up of estimated: 102M - modified BERT-base-Chinese text encoder 26M - 3D U-Net-style vision/anatomy encoder 2.8M - projection layers, anatomy-specific projections, query tokens and…