Keras CAE: Architecture Overview¶
This is the entry point for the Keras CAE documentation. Each of the 10 pipeline steps is covered in a dedicated sub-page. Start here to understand the big picture, then navigate to whichever step you want to understand in depth.
What Is a Convolutional Autoencoder?¶
An autoencoder is a neural network with a single unusual training objective: compress an image into a small internal representation (the bottleneck), then reconstruct the original image from it. There is no external label - the training signal is the difference between the input and the output.
flowchart LR
IMG["Input Image"] --> ENC["Encoder\n(compress)"]
ENC --> BN["Fully Convolutional Bottleneck\n(H/16 x W/16 x C)"]
BN --> DEC["Decoder\n(reconstruct)"]
DEC --> REC["Reconstructed Image"]
IMG & REC --> LOSS["Reconstruction Error\n= Anomaly Score"]
style BN fill:#f0a,color:#fff
style LOSS fill:#f80,color:#fff
Why is this useful for quality control? The network is trained only on defect-free normal images. It learns to reconstruct normal surfaces well. When a defective image is fed to it at test time, the model does not know what the defect looks like - it has never seen one. It tries to reconstruct a normal version of the image and fails. That failure (the high reconstruction error) is the anomaly signal.
Full Pipeline at a Glance¶
flowchart TD
A["MVTec Dataset"] --> B["Load Manifest\nbuild_mvtec_manifest()"]
B --> C{"Category Type?"}
C -- Texture\nwood/leather/carpet --> D["TextureAugmenter\nRotations + Flips + Scale"]
C -- Object\nbottle/transistor/pill --> E["ObjectAugmenter\nColour + Noise only"]
D & E --> F["Otsu + Canny\nForeground Extraction\n(BGRP-G)"]
F --> G["Normalise to 0-1"]
G --> H["Build Keras CAE\nELU + BatchNorm + AdamW"]
H --> I["Train with\nMasked Image Modeling\nSSIM+MSE Loss"]
I --> J["Score Test Images\nTop-K Pooling"]
J --> K["Adaptive Threshold\nQuantile or Mahalanobis"]
K --> L["Evaluate\nAUROC + AUPIMO\n(Train 85/15 Split & Test Set)"]
L --> M{"Anomalies detected?"}
M -- Yes --> N["Generate Reconstruction\nError Heatmap Grid\n(Side-by-side UI)"]
M -- No --> O["Results Dictionary"]
N --> O
style H fill:#4a9,color:#fff
style I fill:#4a9,color:#fff
style L fill:#07a,color:#fff
Hyperparameter Optimization (Optuna)¶
To find the best configuration per category, the pipeline integrates an Optuna study (optuna_study.py).
- Target Metric: Pixel AUPIMO (Area Under the Per-Region Anomaly Detection Curve).
- Search Space: Dynamically adapts based on object vs. texture. For example, Otsu + Canny Foreground Masking is automatically bypassed for textures like Carpet or Wood, preserving structural grammar.
- UI Integration: The Streamlit UI automatically loads optimal values from
data/hyperparameters/keras_cae_best.jsonusing anon_changecallback whenever the category is switched.
Cloud Execution (Google Colab)¶
For large-scale tuning across all 15 MVTec AD categories, you can use the provided Colab notebook:
- Open
notebooks/colab_optuna_sweep.ipynbin Google Colab (or any Jupyter environment). - Execute the notebook to run Optuna studies across all categories. It will iterate through the dataset and write the nested dictionary schema to a file.
- Download the resulting
keras_cae_best.jsonfile from the Colab instance. - Place this file precisely at
data/hyperparameters/keras_cae_best.jsonin your local project root. - Launch the UI (
just run). When you change the category in the Keras CAE tab, it will instantly load and sync these optimal parameters (including both preprocessing checkboxes and model hyperparameters like latent dimensionality).
Module Structure¶
Each Python module handles one specific concern - no module does more than one job:
| Module | What it does |
|---|---|
augmentation.py |
Category-aware data augmentation (texture vs. object strategies) |
segmentation.py |
Otsu + Canny foreground extraction and background replacement |
cae_keras.py |
CAE model definition (build_cae), MIM masking, SSIM+MSE loss, training loop |
scoring.py |
Pixel error map computation, Top-K pooling, adaptive thresholds |
evaluation.py |
AUROC, AUPIMO, Precision/Recall/F1, tradeoff curves |
error_heatmap.py |
Reconstruction Error Heatmap XAI overlays |
cae_pipeline.py |
End-to-end orchestrator: calls all modules in order |
optuna_study.py |
Automated hyperparameter optimization using Optuna (optimizes Pixel AUPIMO) |
Step-by-Step Sub-Pages¶
The 10 pipeline steps are split across focused sub-pages. Each sub-page is self-contained
- you can read any one without reading the others, as long as you understand the overview above.
| Step | What happens | Sub-page |
|---|---|---|
| Step 1 | Category-Aware Data Augmentation | Preprocessing |
| Step 2 | Image Loading, LANCZOS, RGB, BGRP-G, Otsu, Canny | Preprocessing |
| Step 3 | CAE Architecture - Encoder / Decoder layers, ELU, BatchNorm | Model Architecture & Training |
| Step 4 | Masked Image Modeling (MIM) - identity mapping problem | Model Architecture & Training |
| Step 5 | SSIM + MSE Loss - why structure matters more than pixels | Model Architecture & Training |
| Step 6 | AdamW Optimizer - decoupled weight decay, all parameters | Model Architecture & Training |
| Step 7 | Top-K Pooling - noise-robust image-level scoring | Inference & Evaluation |
| Step 8 | Adaptive Threshold - Quantile vs. Mahalanobis | Inference & Evaluation |
| Step 9 | Evaluation Metrics - AUROC, F1, AUPIMO | Inference & Evaluation |
| Step 10 | Heatmap Explainability | Explainability |
Design decision rationale (why not ReLU? why not SAM? why not VAE?) is collected in: Classical Alternatives & Design Decisions
API Endpoint¶
Request body (all fields optional, defaults shown):
{
"data_root": "data/raw/mvtec_ad",
"category": "bottle",
"img_size": 128,
"latent_channels": 32,
"epochs": 20,
"batch_size": 16,
"mask_ratio": 0.25,
"threshold_method": "quantile",
"k_fraction": 0.002,
"use_segmentation": true,
"run_heatmap": true,
"force_retrain": false
}
Response:
{
"status": "success",
"category": "bottle",
"results": {
"auroc": 0.91,
"aupimo": 0.78,
"accuracy": 0.87,
"precision": 0.83,
"recall": 0.90,
"threshold": 0.043217,
"final_train_loss": 0.008431,
"loss_history": [0.124, 0.087, "..."]
}
}
Dependencies¶
All core dependencies are already installed in the Pixi environment. TensorFlow is
installed via the Pixi feature system - no manual pip install is required:
# GPU environment (default): installs tensorflow[and-cuda]
pixi install
# CPU environment: installs tensorflow-cpu (AVX2 + oneDNN)
pixi install --environment ci
For optional SHAP explainability:
For details on how TF chooses GPU vs. CPU at runtime, see TF Device Selection Architecture.
References¶
- Bergmann, P., et al. (2019). Improving Unsupervised Defect Segmentation by Applying Structural Similarity to Autoencoders. VISAPP 2019.
- He, K., et al. (2022). Masked Autoencoders Are Scalable Vision Learners. CVPR 2022.
- Loshchilov, I. & Hutter, F. (2019). Decoupled Weight Decay Regularization. ICLR 2019.
- Clevert, D., Unterthiner, T. & Hochreiter, S. (2016). Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs). ICLR 2016.
- Canny, J. (1986). A Computational Approach to Edge Detection. IEEE TPAMI.
- Otsu, N. (1979). A Threshold Selection Method from Gray-Level Histograms. IEEE SMC.
- Batzner, K., Heckler, L. & Konig, R. (2023). EfficientAD. arXiv 2023.
- Dickson, A. et al. (2024). AUPIMO: Redefining Visual Anomaly Detection Benchmarks with High Speed and Low Tolerance. ECCV 2024.
- Kirillov, A. et al. (2023). Segment Anything. ICCV 2023.
- Lundberg, S. M. & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions (SHAP). NeurIPS 2017.