<!-- .slide: class="cover" -->
<p class="eyebrow">Post-hoc OOD detection · Gradient space</p>

# GROOD

### Gradient-Aware Out-of-Distribution Detection

Mostafa ElAraby · Sabyasachi Sahoo · Yann Pequignot · Paul Novello · Liam Paull

<p class="muted">Transactions on Machine Learning Research</p>

---

## Recognizing unfamiliar inputs

A classifier trained on cats and dogs still produces a prediction for a bird.

**OOD detection adds a rejection decision:** is this input from the distribution of known classes?

- Near-OOD: unfamiliar but semantically similar classes.
- Far-OOD: inputs substantially different from the training distribution.

Note: The presentation focuses on semantic OOD. Covariate shifts with the same labels are a related but distinct evaluation problem.

---

## Signals for OOD detection

| Signal | Typical approach | Challenge |
| :--- | :--- | :--- |
| Output | Softmax confidence, energy | Confident predictions on unknowns |
| Features | Distance to ID examples | Near-OOD may resemble known classes |
| Gradients | Sensitivity of a loss | Which parameters and which reference? |

**GROOD measures sensitivity to an artificial OOD prototype.**

---

## Prototype geometry

- Neural collapse motivates representing each ID class by its mean feature.
- ID features tend to cluster near their class prototypes.
- An artificial OOD prototype provides a shared reference for sensitivity.

> Measure how an input responds to the reference, then compare that response with familiar inputs.

Note: Neural collapse motivates the method; it is not a guarantee that every architecture and checkpoint has perfectly collapsed features.

---

## GROOD pipeline

![GROOD pipeline: input, feature extractor, features, distances to prototypes, and gradient-space OOD score](assets/grood-slide8-0.webp)

**Features → prototype distances → gradient vector → nearest ID gradient**

<p class="caption">Method illustration from the supplied GROOD presentation</p>

---

## ID class prototypes

For each class $c$, average its training features:

<div class="equation">$$p_c^l=\frac1{|\mathcal D_c|}\sum_{x\in\mathcal D_c} f^l(x)$$</div>

Compute prototypes at the **early** and **penultimate** layers.

- Early prototypes support synthetic feature generation.
- Penultimate prototypes define the distance-based classifier.

---

## Constructing the OOD prototype

Interpolate early features toward the second-closest class prototype:

<div class="equation">$$\hat h(x)=f^{\mathrm{mid}}\!\left(\lambda f^{\mathrm{early}}(x)+(1-\lambda)p_{c_2}^{\mathrm{early}}\right)$$</div>

Average the synthetic penultimate features to obtain $p_{\mathrm{ood}}$.

**The synthetic variant uses ID data to create a reference without real OOD examples.**

Note: This is post-hoc feature interpolation, not a requirement to retrain the backbone with mixup. Other prototype construction variants use auxiliary data; they should be evaluated separately.

---

## Distance-based probabilities

Use **negative Euclidean distances** as logits:

<div class="equation">$$L_i(h)=-\|h-p_i\|_2$$</div>

Include the $C$ class prototypes and the artificial OOD prototype.

<div class="equation">$$q_i(h)=\frac{\exp L_i(h)}{\sum_{j=1}^{C+1}\exp L_j(h)}$$</div>

<p class="source"><a href="https://arxiv.org/html/2312.14427#S4.SS1">Paper §4.1, Eqs. 2–3</a></p>

Note: The supplied slide deck and blog use squared distances alongside a normalized-direction gradient. The published paper uses unsquared L2 distances, which yield the gradient shown on the next slide. q denotes a probability; p denotes a prototype.

---

## Gradient with respect to the reference

For cross-entropy with any ID class label:

<div class="equation">$$g(h)=\nabla_{p_{\mathrm{ood}}}H(h,y)=q_{\mathrm{ood}}(h)\frac{h-p_{\mathrm{ood}}}{\|h-p_{\mathrm{ood}}\|_2}$$</div>

- **Magnitude:** $q_{\mathrm{ood}}(h)$.
- **Direction:** from the reference toward the input’s feature.
- Closed-form computation avoids backpropagation through the backbone.

<p class="source"><a href="https://arxiv.org/html/2312.14427#A1.SS6">Paper §4.1 and Appendix A.6</a></p>

Note: y is an ID label, so the OOD component of the cross-entropy derivative is q_ood. The main paper has an inconsistent C+1 label notation; the expression is consistent with an ID target and unsquared negative L2 logits. Stabilize the denominator for a zero-distance input.

---

## Gradient-space geometry

![t-SNE projection of ID classes and OOD samples in gradient space](assets/grood-slide12-0.webp)

ID inputs yield coherent gradient responses; unfamiliar inputs can differ in magnitude and direction.

<p class="caption">Qualitative projection from the supplied presentation; t-SNE alone does not establish detection accuracy.</p>

---

## Nearest-neighbor OOD score

Store gradient vectors from ID training inputs.

<div class="equation">$$S(x)=\min_{x_i\in\mathcal D_{\mathrm{in}}}\|g(h(x))-g(h(x_i))\|_2$$</div>

**Low score:** response resembles a familiar input.

**High score:** response is unusual → more OOD-like.

Note: This score has the opposite orientation to TOOD's ID score. For a threshold tau, accept ID if S is at most tau and reject otherwise.

---

## Setup and inference

**Setup on a frozen model**

1. Compute class prototypes and the OOD reference.
2. Compute the ID gradient bank.
3. Build the nearest-neighbor search index.

**At inference**

Extract features → compute the closed-form gradient → query the index.

<p class="muted">Post-hoc detection; nearest-neighbor search adds storage and inference cost.</p>

Note: Search can use FAISS. The detector leaves the original classifier's class predictions intact and computes a separate rejection score.

---

## Experimental protocol

- OpenOOD v1.5 evaluation on CIFAR-10, CIFAR-100, ImageNet-200, ImageNet-1K.
- Separate Near-OOD and Far-OOD results.
- ResNet backbones, plus ViT and Swin-T evaluations.
- Comparisons include MSP, ODIN, ReAct, KNN, ViM and NCI.

**AUROC (%) measures separation across thresholds; higher is better.**

---

## Near-OOD results

AUROC (%) · OpenOOD v1.5

| Method | CIFAR-10 | CIFAR-100 | IN-200 | IN-1K |
| :--- | ---: | ---: | ---: | ---: |
| MSP | 88.0 | 80.3 | 83.3 | 76.0 |
| ReAct | 87.1 | 80.7 | 81.9 | 77.4 |
| KNN | 90.6 | 80.2 | 81.6 | 71.1 |
| ViM | 88.7 | 75.0 | 78.7 | 72.1 |
| NCI | 88.8 | **81.0** | **83.5** | 78.6 |
| GROOD | **91.2** | 78.9 | 83.4 | **78.9** |

**Strong CIFAR-10 and ImageNet-1K results; performance varies by dataset.**

Note: Selected baseline rows from the provided presentation/blog and official repository. Bold marks the best value among the displayed methods, not a claim about every published method.

---

## Far-OOD results

AUROC (%) · OpenOOD v1.5

| Method | CIFAR-10 | CIFAR-100 | IN-200 | IN-1K |
| :--- | ---: | ---: | ---: | ---: |
| MSP | 90.7 | 77.8 | 90.1 | 85.2 |
| ReAct | 90.4 | 80.4 | 92.3 | 93.7 |
| KNN | 93.0 | 82.4 | 93.2 | 90.2 |
| ViM | 93.5 | 81.7 | 91.3 | 92.7 |
| NCI | 91.3 | 81.3 | **93.7** | **95.5** |
| GROOD | **93.8** | **84.4** | 92.2 | 94.8 |

**GROOD leads the displayed methods on CIFAR-10 and CIFAR-100.**

---

## Checkpoint robustness

![AUROC across model checkpoints for multiple OOD detectors](assets/grood-slide25-0.webp)

<p class="caption">Checkpoint comparison from the supplied GROOD presentation</p>

Detection quality should be assessed alongside the choice of classifier checkpoint.

Note: Read this as an empirical comparison on the plotted datasets and checkpoints, not a guarantee of stability for all training procedures.

---

## Scope and limitations

- Requires access to ID training features and intermediate model representations.
- Prototype construction and backbone geometry affect detection quality.
- A single OOD reference does not represent every possible unknown class.
- The gradient bank and nearest-neighbor index add memory and search cost.

**Compare full gradient vectors, rather than relying only on confidence or gradient norm.**

---

<!-- .slide: class="cover" -->
<p class="eyebrow">Questions & discussion</p>

# GROOD

[Paper · OpenReview](https://openreview.net/forum?id=2V7itvvMVJ)

[Paper · arXiv:2312.14427](https://arxiv.org/abs/2312.14427)

[Code · Gradient-Aware-OOD-Detection](https://github.com/mostafaelaraby/Gradient-Aware-OOD-Detection)

[Related presentation · TOOD](TOOD.html)

Note: Source material: TMLR- GROOD.pptx and the site's GROOD publication post. Distance and gradient conventions were checked against the published paper, version 3.
