Last released Dec 8, 2024
A virtual environment that generates vision-language tasks with varying complexity.
Supported by