<!-- .slide: class="cover" -->
<p class="eyebrow">Continual learning · OOD detection</p>

# TOOD

### Task-Aware Out-of-Distribution Score Calibration for Continual Learners

Mostafa ElAraby · Samer B. Nashed · Liam Paull

<p class="muted">CoLLAs 2026 · Oral presentation</p>

Note: The presentation focuses on why classification retention alone does not establish reliable OOD detection.

---

## OOD detection throughout learning

- A continual learner acquires new classes over time.
- It must still recognize familiar inputs from earlier tasks.
- It must also reject inputs outside **all classes learned so far**.

> Remembering class labels does not guarantee remembering the boundary to the unknown.

Note: Example: cups and plates are learned first, books and pens next. A shoe remains OOD throughout. Books become ID after their task is learned.

---

## Evaluation setting

**Class-incremental learning:** disjoint class groups arrive in successive tasks.

<div class="equation">$$\mathcal C_{\mathrm{seen},t}=\bigcup_{i=0}^{t}\mathcal C_i$$</div>

- Classify without knowing the test input’s task.
- Evaluate OOD detection after every task.
- Keep evaluation OOD classes outside the learned classes throughout the stream.

Note: This fixed OOD evaluation isolates deterioration in detection from the changing definition of which classes are known.

---

## Measuring OOD forgetting

With tasks indexed $0,\ldots,N-1$, let $R_{t,i}$ be OOD AUROC for task $i$ at checkpoint $t$.

<div class="equation">$$D_i=R_{i,i}-R_{N-1,i}$$</div>

Average incremental AUROC tracks the full learning trajectory:

<div class="equation">$$\overline R=\frac1N\sum_{t=0}^{N-1}\frac1{t+1}\sum_{i=0}^{t}R_{t,i}$$</div>

<p class="muted">Higher AUROC is better. Positive $D_i$ indicates deterioration.</p>

Note: AUROC comparisons use each task's ID examples against the evaluation OOD sets. The trajectory metric is different from final-checkpoint AUROC.

---

## Classification and OOD forgetting

![Diagnostic comparison of representation change, OOD forgetting, and classification forgetting](assets/tood-slide12-0.webp)

<p class="caption">Diagnostic plots from the supplied TOOD presentation</p>

**Classification performance alone is an incomplete reliability measure.**

Note: Read the original figure axes carefully: the diagnostic compares representation change with two forms of forgetting. It motivates examining output scales as well as features.

---

## Mechanism 1: the confidence gap

![Task-zero ID energy falls toward the OOD energy band as tasks are learned](assets/tood-slide13-0.webp)

- Older tasks’ logit scales weaken relative to newer tasks.
- Correct class ordering can survive while energy separation deteriorates.

Note: This is the mechanism TOOD targets. The diagnostic follows an old task over a CIFAR-10 learning sequence; it is not an average across all tasks.

---

## Mechanism 2: manifold crowding

![Feature-space diagnostic illustrating crowding as additional tasks are learned](assets/tood-slide14-0.webp)

New classes occupy feature-space regions that previously separated known inputs from outliers.

**TOOD corrects output scores; it does not restore lost feature-space margins.**

---

<!-- .slide: class="animation-slide" -->
## TOOD: animated overview

<img data-src="assets/tood-calibration.gif" width="1200" height="675" alt="Animated illustration of old-task ID scores drifting toward OOD scores, followed by TOOD grouping task responses, aligning their scores, and taking the strongest response">

<p class="caption">Illustration of score drift and calibration, not measured experimental data</p>

Note: Animation supplied from LinkedIn. It illustrates the core maximum-score variant without the optional top-two margin. Calibration uses familiar examples; the input's task label is not required at prediction time.

---

## Why a global score shift is insufficient

AUROC depends on sample ranking.

<div class="equation">$$\operatorname{AUROC}(g\circ S)=\operatorname{AUROC}(S)$$</div>

for any **strictly increasing** transformation $g$.

TOOD calibrates each task’s response **before combining responses**, allowing different samples to change their order.

Note: Strictly increasing is necessary here: a decreasing transformation reverses score orientation, while a non-strict transformation can introduce ties.

---

## Step 1: per-task energy

Group the classifier’s logits by the classes introduced in each task:

<div class="equation">$$E_t(x)=\log\sum_{c\in\mathcal C_t}\exp(h_c(x))$$</div>

- One response for every task seen so far.
- Preserve information obscured by a single global score.
- Use the sign-reversed energy convention: **higher means more ID-like**.

Note: Task partitions are known from training. The input's task identity is not supplied at inference.

---

## Step 2: mean-shift calibration

Estimate each task’s mean energy $\mu_t$ on its own familiar calibration examples.

<div class="equation">$$E_t^{\mathrm{ms}}(x)=E_t(x)+(\mu_{\mathrm{ref}}-\mu_t)$$</div>

The newest task provides the reference.

**Align the location of task scores while preserving their spread.**

Note: Recompute statistics for all seen tasks under the current frozen model after each new task. Replay memory can supply the examples; otherwise use a held-out ID memory.

---

## Step 2: robust-anchor calibration

Use each task’s median $\tilde e_t$ and scaled median absolute deviation $\widehat{\mathrm{MAD}}_t$:

<div class="equation">$$E_t^{\mathrm{rob}}(x)=\frac{E_t(x)-\tilde e_t}{\widehat{\mathrm{MAD}}_t}\widehat{\mathrm{MAD}}_{\mathrm{ref}}+\tilde e_{\mathrm{ref}}$$</div>

**Align location and spread using statistics less sensitive to outliers.**

<p class="muted">Requires representative ID memory and a nonzero, numerically stabilized scale.</p>

---

## Step 3: combine calibrated responses

Core score:

<div class="equation">$$S_{\mathrm{TOOD}}(x)=\max_t E_t^{\mathrm{norm}}(x)$$</div>

Optional top-two margin, with $E_{(1)}\ge E_{(2)}$:

<div class="equation">$$S_\lambda(x)=E_{(1)}^{\mathrm{norm}}+\lambda\left(E_{(1)}^{\mathrm{norm}}-E_{(2)}^{\mathrm{norm}}\right)$$</div>

<p class="muted">$\lambda=0$ gives the core score; main experiments use $\lambda=0.5$.</p>

Note: A high response from at least one task suggests ID. The margin rewards a clear task preference. With one task, use the core score rather than a nonexistent runner-up.

---

## Calibration and inference

**After each task**

1. Freeze the current learner.
2. Forward the ID memory through it.
3. Estimate and store statistics for every seen task.

**For each test input**

Compute task energies → normalize → maximum plus optional margin.

<p class="muted">No retraining, no OOD calibration examples, no change to class predictions.</p>

---

## Experimental protocol

| Dataset | Learning stream | Backbone |
| :--- | :--- | :--- |
| CIFAR-10 | 5 × 2 classes | ResNet-32 |
| CIFAR-100 | 10 × 10 classes | ResNet-32 |
| ImageNet-1K | 100 × 10 classes | ResNet-18 |

- CIFAR learners: iCaRL, BiC, DER, WA, LwF.
- Near- and Far-OOD evaluation following OpenOOD.
- ID memory: 200 examples for CIFAR-10; 700 for CIFAR-100.

Note: Main CIFAR results average three seeds. The ImageNet scalability study uses one seed. ViT-B/16 appears in additional diagnostics.

---

## CIFAR-10 results

Average incremental AUROC (%) · mean ± standard deviation, three seeds

| Learner | Energy | Mean Shift | Robust Anchor |
| :--- | ---: | ---: | ---: |
| iCaRL | 65.9 ± 1.5 | 66.5 ± 1.5 | 66.4 ± 1.4 |
| BiC | 67.2 ± 1.0 | **71.6 ± 0.2** | 70.6 ± 0.3 |
| DER | 61.2 ± 0.7 | **69.3 ± 0.9** | 68.6 ± 0.9 |
| WA | 78.1 ± 2.3 | 78.3 ± 2.1 | 78.2 ± 2.4 |
| LwF | 72.2 ± 0.5 | 73.2 ± 0.3 | 73.0 ± 0.7 |

**DER: +8.1 AUROC points over uncalibrated energy.**

Note: Results are averaged over the learning sequence and Near-/Far-OOD evaluation, as reported in the local blog post and supplied deck.

---

## CIFAR-100 results

Average incremental AUROC (%) · mean ± standard deviation, three seeds

| Learner | Energy | Mean Shift | Robust Anchor |
| :--- | ---: | ---: | ---: |
| iCaRL | 58.1 ± 1.3 | 58.4 ± 1.0 | 58.9 ± 1.0 |
| BiC | 65.6 ± 0.5 | 68.4 ± 1.3 | **68.8 ± 1.6** |
| DER | 64.6 ± 1.2 | **67.1 ± 0.5** | 66.1 ± 0.9 |
| WA | **71.9 ± 0.1** | 71.7 ± 0.2 | 71.5 ± 0.1 |
| LwF | **67.0 ± 1.6** | 66.7 ± 1.7 | 66.3 ± 1.6 |

**Benefit depends on the learner; WA and LwF slightly decline.**

---

## ImageNet-1K: 100 tasks

Average incremental AUROC (%) · single seed

| Learner | Energy | Mean Shift | Robust Anchor |
| :--- | ---: | ---: | ---: |
| BiC | 59.2 | **65.2** | 65.1 |
| WA | **70.1** | 65.7 | 65.6 |
| DER | 70.2 | **71.4** | 70.5 |

Mean Shift improves BiC by **6.0 points** and DER by **1.2 points**.

WA falls by **4.4 points**: calibration is not a universal upgrade.

---

## Scope and limitations

- Requires known training task partitions and representative ID memory.
- Corrects task-wise score drift, rather than feature-space crowding.
- Gains vary with the learner and its existing output correction.
- Evaluate detection throughout learning, alongside classification retention.

> Calibrate the response to each known task before deciding whether an input is unknown.

---

<!-- .slide: class="cover" -->
<p class="eyebrow">Questions & discussion</p>

# TOOD

[Paper · arXiv:2607.29592](https://arxiv.org/abs/2607.29592)

[Code · tood-continual-ood](https://github.com/mostafaelaraby/tood-continual-ood)

[Related presentation · GROOD](GROOD.html)

<p class="muted">Mostafa ElAraby · Mila / Université de Montréal</p>

Note: Source material: TOOD paper presentation.pptx and the site's TOOD publication post. Current paper metadata is verified against arXiv. Conflicting top-two ranking counts in the supplied materials are omitted; numerical results are shown directly.
