Overview of the Image Classification Benchmark
mainThe Image Classification Benchmark evaluates high-performance data processing for multimodal workloads. It classifies 803,580 rows (consisting of 80,358 unique images, each repeated 10x) using the ResNet18 model. The workflow involves downloading images, applying preprocessing transforms, and running inference to predict ImageNet labels across distributed GPU nodes.
Benchmark Specifications:
- Input Dataset: ImageNet benchmark dataset in S3 parquet format.
- Output Format: Parquet files containing image URLs and predicted labels.
- Infrastructure: 8 worker nodes using
g6.xlargeinstances. - Framework Versions used in benchmark: Daft 0.6.2, Ray Data 2.49.2, AWS EMR Spark 7.10.0.