Command Line Interface (CLI)
The AutoMIL Command Line Interface (CLI) provides a user-friendly way to interact with the framework directly from the terminal. It allows users to perform various tasks such as data preprocessing, model training, evaluation, and prediction without needing to write any code.
run_pipeline
run_pipeline(
slide_dir: Path,
annotation_file: Path,
project_dir: Path,
patient_column: str,
label_column: str,
slide_column: str | None,
resolutions: str,
model: str,
k: int,
split_file: str | None,
tissue_detection: str,
stain_normalizer: str,
transform_labels: bool,
is_pretiled: bool,
verbose: bool,
)
Execute the complete AutoMIL pipeline for whole slide image analysis.
This command runs the full AutoMIL workflow, including project setup, dataset preparation, model training with k-fold cross-validation, evaluation, and result visualization.
Pipeline stages:
- Project setup and configuration
- Dataset preparation and tile extraction
- Model training with k-fold cross-validation
- Model evaluation and ensemble creation
- Result visualization
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
slide_dir
|
str | Path
|
Directory containing whole-slide images or pre-extracted tiles. |
required |
annotation_file
|
str | Path
|
CSV file containing slide- or patient-level annotations and labels. |
required |
project_dir
|
str | Path
|
Output directory where trained models and intermediate files will be written. |
required |
patient_column
|
str
|
Name of the column containing patient identifiers. |
required |
label_column
|
str
|
Name of the column containing class labels. |
required |
slide_column
|
str | None
|
Name of the column containing slide identifiers. |
required |
resolutions
|
str
|
Comma-separated list of resolution presets to train on. |
required |
model
|
str
|
Model architecture to train. |
required |
k
|
int
|
Number of folds used for k-fold cross-validation. |
required |
is_pretiled
|
bool
|
Indicates that the input slides are already tiled. |
required |
transform_labels
|
bool
|
If enabled, transforms labels to floating-point values. |
required |
verbose
|
bool
|
Enables verbose logging output. |
required |
Examples
Basic usage with default settings:
automil run-pipeline /data/slides /data/annotations.csv ./results
Multi-resolution training with verbose output:
automil run-pipeline -r "Low,High" -v /data/slides /data/annotations.csv ./results
Custom model and k-fold settings:
automil run-pipeline -m TransMIL -k 5 /data/slides /data/annotations.csv ./results
Skip tiling if tiles are pre-extracted:
automil run-pipeline -p /data/slides /data/annotations.csv ./results
Custom column names in the annotation file:
automil run-pipeline -pc "patient_name" -lc "diagnosis" -sc "slide_name" /data/slides /data/annotations.csv ./results
Provide a predefined train-test split:
automil run-pipeline --split-file /data/split.json /data/slides /data/annotations.csv ./results
Annotation file requirements
The annotation file must be a CSV file containing at least the following columns:
- Patient identifiers (default column name:
patient) - Slide identifiers (default column name:
slide; optional) - Class labels (default column name:
label)
By default, AutoMIL looks for columns named patient, slide, and label.
These defaults can be overridden using the --patient_column,
--slide_column, and --label_column options.
Minimal annotation file example
patient,slide,label
001,001_1,0
001,001_2,0
002,002,1
003,003,1
Expected slide directory structure
SLIDE_DIR should contain whole slide images in supported formats
such as .svs, .tiff, or .png.
Example structure:
/data/slides/
|-- slide1.svs
|-- slide2.tiff
|-- slide3.tiff
PNG Slide Handling
If slides are in PNG, AutoMIL will first convert them to TIFF for easier processing.
Using pretiled data
If tiles have already been extracted from the slides, use the --is_pretiled flag.
In the case of pretiled data, AutoMIL expects the following directory structure for SLIDE_DIR:
/data/slides/
|-- slide1/
| |-- tile_0_0.png
| |-- tile_0_1.png
| |-- ...
|-- slide2/
| |-- tile_0_0.png
| |-- tile_0_1.png
| |-- ...
Slide name matching
Tile names are arbitrary but slide subdirectories must match the slide names in ANNOTATION_FILE.
Providing a train-test split
Use the --split-file option to provide a JSON file defining train-test splits.
The JSON file will have the following structure:
{
"train": ["slide1", "slide2", ...],
"test": ["slide3", "slide4", ...]
}
or:
{
"train": ["slide1", "slide2", ...],
"validation": ["slide3", "slide4", ...]
}
Output structure
project_dir/
├── bags/ # Extracted tile features
├── models/ # Trained model checkpoints
├── ensemble/ # Ensemble predictions
├── annotations.csv # Processed annotations
└── results.json # Performance metrics
Source code in automil/cli/commands/run_pipeline.py
40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 | |
train
train(
slide_dir: Path,
annotation_file: Path,
project_dir: Path,
patient_column: str,
label_column: str,
slide_column: str | None,
resolutions: str,
model: str,
tissue_detection: str,
stain_normalizer: str,
k: int,
is_pretiled: bool,
transform_labels: bool,
verbose: bool,
)
Train one or more MIL models on a given dataset.
This command initializes an AutoMIL project, prepares the dataset, and trains MIL models using k-fold cross-validation. Training can be performed at one or multiple resolution presets.
Pipeline stages:
- Project setup and configuration
- Dataset preparation and tile extraction
- Model training with k-fold cross-validation
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
slide_dir
|
str | Path
|
Directory containing whole-slide images or pre-extracted tiles. |
required |
annotation_file
|
str | Path
|
CSV file containing slide- or patient-level annotations and labels. |
required |
project_dir
|
str | Path
|
Output directory where trained models and intermediate files will be written. |
required |
patient_column
|
str
|
Name of the column containing patient identifiers. |
required |
label_column
|
str
|
Name of the column containing class labels. |
required |
slide_column
|
str | None
|
Name of the column containing slide identifiers. |
required |
resolutions
|
str
|
Comma-separated list of resolution presets to train on. |
required |
model
|
str
|
Model architecture to train. |
required |
k
|
int
|
Number of folds used for k-fold cross-validation. |
required |
is_pretiled
|
bool
|
Indicates that the input slides are already tiled. |
required |
transform_labels
|
bool
|
If enabled, transforms labels to floating-point values. |
required |
verbose
|
bool
|
Enables verbose logging output. |
required |
Examples
Basic usage with default settings:
automil train /data/slides /data/annotations.csv ./results
Multi-resolution training with verbose output::
automil train -r "Low,High" -v /data/slides /data/annotations.csv ./results
Custom model and 5-fold configuration:
automil train -m TransMIL -k 5 /data/slides /data/annotations.csv ./results
Using pre-tiled slides::
automil train -p /data/slides /data/annotations.csv ./results
Annotation file requirements
The annotation file must be a CSV file containing at least the following columns:
- Patient identifiers (default column name:
patient) - Slide identifiers (default column name:
slide; optional) - Class labels (default column name:
label)
By default, AutoMIL looks for columns named patient, slide, and label.
These defaults can be overridden using the --patient_column,
--slide_column, and --label_column options.
Minimal annotation file example
patient,slide,label
001,001_1,0
001,001_2,0
002,002,1
003,003,1
Expected slide directory structure
SLIDE_DIR should contain whole slide images in supported formats
such as .svs, .tiff, or .png.
Example structure:
/data/slides/
|-- slide1.svs
|-- slide2.tiff
|-- slide3.tiff
PNG Slide Handling
If slides are in PNG, AutoMIL will first convert them to TIFF for easier processing.
Using pretiled data
If tiles have already been extracted from the slides, use the --is_pretiled flag.
In the case of pretiled data, AutoMIL expects the following directory structure for SLIDE_DIR:
/data/slides/
|-- slide1/
| |-- tile_0_0.png
| |-- tile_0_1.png
| |-- ...
|-- slide2/
| |-- tile_0_0.png
| |-- tile_0_1.png
| |-- ...
Slide name matching
Tile names are arbitrary but slide subdirectories must match the slide names in ANNOTATION_FILE.
Providing a train-test split
Use the --split-file option to provide a JSON file defining train-test splits.
The JSON file will have the following structure:
{
"train": ["slide1", "slide2", ...],
"test": ["slide3", "slide4", ...]
}
or:
{
"train": ["slide1", "slide2", ...],
"validation": ["slide3", "slide4", ...]
}
Output structure
project_dir/
├── bags/ # Extracted tile features
├── models/ # Trained model checkpoints
├── ensemble/ # Ensemble predictions
├── annotations.csv # Processed annotations
└── results.json # Performance metrics
Source code in automil/cli/commands/train.py
39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 | |
predict
predict(
slide_dir: str | Path,
annotation_file: str | Path,
bags_dir: str | Path,
model_dir: str | Path,
output_dir: str | Path,
patient_column: str,
label_column: str,
slide_column: str | None,
verbose: bool,
)
Generate predictions using one or more trained MIL models.
This command loads trained model checkpoints and generates predictions
for the slides in SLIDE_DIR using precomputed tile feature bags from
BAGS_DIR. Predictions are written to the specified output directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
slide_dir
|
str | Path
|
Directory containing whole-slide images. |
required |
annotation_file
|
str | Path
|
CSV file containing slide- or patient-level annotations and labels. |
required |
bags_dir
|
str | Path
|
Directory containing extracted tile feature bags. |
required |
model_dir
|
str | Path
|
Directory containing trained model checkpoints. |
required |
output_dir
|
str | Path
|
Directory to which prediction files will be written. |
required |
patient_column
|
str
|
Name of the column containing patient identifiers. |
required |
label_column
|
str
|
Name of the column containing class labels. |
required |
slide_column
|
str | None
|
Name of the column containing slide identifiers. |
required |
verbose
|
bool
|
Enables verbose logging output. |
required |
Examples
Basic usage with multiple models:
automil predict /data/slides /data/annotations.csv /data/bags /data/models -o ./predictions
Generate predictions with a single model:
automil predict /data/slides /data/annotations.csv /data/bags /data/models/model_1 -v
Override annotation column names:
automil predict -pc "patient_id" -lc "outcome" -sc "slide_id" /data/slides /data/annotations.csv /data/bags /data/models/model_1 -o ./predictions
Expected model directory structure
MODEL_DIR may either point to a single model directory or to a parent
directory containing multiple model subdirectories.
Single model example:
/data/models/model_1/
|-- best_valid.pth
|-- ...
Multiple models example:
/data/models/
|-- model_1/
| |-- best_valid.pth
|-- model_2/
| |-- best_valid.pth
| |-- ...
Multiple models
When multiple models are provided, AutoMIL generates a separate prediction file for each model.
Annotation file requirements
The annotation file must be a CSV file containing at least the following columns:
- Patient identifiers (default column name:
patient) - Slide identifiers (default column name:
slide; optional) - Class labels (default column name:
label)
By default, AutoMIL looks for columns named patient, slide, and label.
These defaults can be overridden using the --patient_column,
--slide_column, and --label_column options.
Minimal annotation file example
patient,slide,label
001,001_1,0
001,001_2,0
002,002,1
003,003,1
Output directory format
OUTPUT_DIR must be a directory path. Prediction results are saved as
separate .csv or .parquet files inside this directory.
When multiple models are used, output files include a suffix indicating the corresponding model.
Source code in automil/cli/commands/predict.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 | |
evaluate
evaluate(
slide_dir: str | Path,
annotation_file: str | Path,
bags_dir: str | Path,
model_dir: str | Path,
output_dir: str | Path,
patient_column: str,
label_column: str,
slide_column: str | None,
verbose: bool,
)
Evaluate one or more trained MIL models on a labeled dataset.
This command generates predictions for the slides in SLIDE_DIR using
trained models from MODEL_DIR and corresponding tile feature bags from
BAGS_DIR. The predictions are evaluated against the provided annotations,
and summary metrics and plots are generated.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
slide_dir
|
str | Path
|
Directory containing whole-slide images. |
required |
annotation_file
|
str | Path
|
CSV file containing slide- or patient-level annotations and labels. |
required |
bags_dir
|
str | Path
|
Directory containing extracted tile feature bags. |
required |
model_dir
|
str | Path
|
Directory containing trained model checkpoints. |
required |
output_dir
|
str | Path
|
Directory to which evaluation results will be written. |
required |
patient_column
|
str
|
Name of the column containing patient identifiers. |
required |
label_column
|
str
|
Name of the column containing class labels. |
required |
slide_column
|
str | None
|
Name of the column containing slide identifiers. |
required |
verbose
|
bool
|
Enables verbose logging output. |
required |
Examples
Evaluate a single model:
automil evaluate /data/slides /data/annotations.csv /data/bags /data/models/model_1 -o ./results
Evaluate multiple models:
automil evaluate /data/slides /data/annotations.csv /data/bags /data/models -v
Override annotation column names:
automil evaluate -pc "patient_id" -lc "outcome" -sc "slide_id" /data/slides /data/annotations.csv /data/bags /data/models/model_1 -o ./results
Expected model directory structure
MODEL_DIR may refer either to a single model directory or to a parent
directory containing multiple model subdirectories.
Single model example:
/data/models/model_1/
|-- best_valid.pth
|-- ...
Multiple models example:
/data/models/
|-- model_1/
| |-- best_valid.pth
|-- model_2/
| |-- best_valid.pth
| |-- ...
Multiple models
When multiple models are evaluated, AutoMIL generates separate evaluation results for each model and compares their performance.
Annotation file requirements
The annotation file must be a CSV file containing at least the following columns:
- Patient identifiers (default column name:
patient) - Slide identifiers (default column name:
slide; optional) - Class labels (default column name:
label)
By default, AutoMIL looks for columns named patient, slide, and label.
These defaults can be overridden using the --patient_column,
--slide_column, and --label_column options.
Minimal annotation file example
patient,slide,label
001,001_1,0
001,001_2,0
002,002,1
003,003,1
Output directory format
OUTPUT_DIR must be a directory path. Evaluation results, metrics, and plots
are written to this directory.
When multiple models are evaluated, output files include a suffix indicating the corresponding model.
Source code in automil/cli/commands/evaluate.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 | |
create_split
create_split(
slide_dir: str | Path,
annotation_file: str | Path,
output_file: str | Path,
test_fraction: float,
read_only: bool,
verbose: bool,
)
Create a train–test split file from dataset annotations.
This command reads the provided annotation file and generates a train–test split, which is saved as a JSON file. The resulting split can be reused for reproducible training and evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
slide_dir
|
str | Path
|
Directory containing whole-slide images. |
required |
annotation_file
|
str | Path
|
CSV file containing slide- or patient-level annotations and labels. |
required |
output_file
|
str | Path
|
Path to which the split JSON file will be written. |
required |
test_fraction
|
float
|
Fraction of samples to assign to the test set. |
required |
read_only
|
bool
|
If enabled, an existing split file will not be overwritten. |
required |
verbose
|
bool
|
Enables verbose logging output. |
required |
Examples
Create a split with default settings:
automil create-split /data/slides /data/annotations.csv -o split.json
Create a split without overwriting an existing file:
automil create-split /data/slides /data/annotations.csv -o split.json --read-only
Output file format
The output JSON file contains slide identifiers grouped by split name.
Example structure:
{
"train": ["slide1", "slide2", ...],
"test": ["slide3", "slide4", ...]
}
Depending on the configuration, a validation split may be generated
instead of or in addition to a test split.
Source code in automil/cli/commands/create_split.py
31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | |