Tiny Vision Foundry

Tiny Vision Foundry compares a compact convolutional teacher with two copies of the same tiny student architecture:

  1. a student trained from labels alone;
  2. a student trained from labels plus the teacher's softened predictions.

The experiment uses the 8x8 handwritten-digit corpus bundled with scikit-learn. The data recipe converts it into deterministic train, validation, and test Parquet splits that load directly through Hugging Face Datasets.

Reproduce

uv run python projects/tiny-vision-foundry/prepare_data.py
uv run python projects/tiny-vision-foundry/train.py

All three networks use the same split and preprocessing. Final test metrics are computed only after training and model selection are complete.

Verified results

Variant Parameters Test accuracy Macro F1
Teacher CNN 13,706 97.78% 0.9771
Tiny student, labels only 2,198 98.52% 0.9848
Tiny student, distilled 2,198 97.41% 0.9730

The tiny student is 6.24 times smaller than the teacher. In this first experiment, knowledge distillation reduced accuracy by 1.11 percentage points versus the identical student trained from labels alone. That result is preserved as a negative ablation; the labels-only student is the recommended artifact.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support