Add model card metadata and description
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -1,3 +1,30 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: apache-2.0
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: transformers
|
| 4 |
+
pipeline_tag: text-classification
|
| 5 |
+
---
|
| 6 |
+
|
| 7 |
+
# ShadowMem Judge Model
|
| 8 |
+
|
| 9 |
+
This repository contains a judge model from the paper [Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory](https://huggingface.co/papers/2605.03228).
|
| 10 |
+
|
| 11 |
+
ShadowMem is a defensive framework that maintains a dedicated, safety-focused agentic memory — inspired by the *shadow stack* abstraction in systems security — to distill and retain safety-critical context across an agent's full execution trajectory, and uses this shadow memory to proactively assess pending actions before they are executed.
|
| 12 |
+
|
| 13 |
+
The model is the trained shadow-memory judge used to evaluate pending actions during agent execution.
|
| 14 |
+
|
| 15 |
+
Code: https://github.com/ZJUWYH/ShadowMem
|
| 16 |
+
|
| 17 |
+
## Citation
|
| 18 |
+
|
| 19 |
+
```bibtex
|
| 20 |
+
@inproceedings{wang2026shadowmem,
|
| 21 |
+
author = {Yuhui Wang and Tanqiu Jiang and Jiacheng Liang and Charles Fleming and Ting Wang},
|
| 22 |
+
title = {Safeguarding {LLM} Agents against Long-Horizon Threats via Shadow Memory},
|
| 23 |
+
booktitle = {Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security},
|
| 24 |
+
series = {CCS '26},
|
| 25 |
+
year = {2026},
|
| 26 |
+
publisher = {Association for Computing Machinery},
|
| 27 |
+
doi = {10.1145/3830454.3846601},
|
| 28 |
+
url = {https://doi.org/10.1145/3830454.3846601}
|
| 29 |
+
}
|
| 30 |
+
```
|