目录 / Test-robustness-ai
SKILL
未评级
已上架
Test-robustness-ai
## 🔬 Adversarial Testing on Abliterated Models This section covers testing AI models that have been "uncensored" (abliterated). 1. **Load the Model**: Use `transformers` to load the abliterated model you want to test. 2. **Run a "Forbidden" Prompt Set**: Use a dataset of prompts that would normally be refused (e.g., HarmBench, AdvBench). 3. **Evaluate Responses**: - **Refusal Rate**: How many prompts were refused? Target should be < 5%. - **Coherence & Quality**: Does the model maintain consistent responses, or does it hallucinate? 4. **Generate a Report**: Produce a report comparing original and abliterated model performance.
存档时间线
| 版本 | 存档时间 | 内容哈希 | 内容 |
|---|---|---|---|
| v1 | 2026-09-28 01:48 | 16dd6768 | 可取 |
版本索引永久保留;内容副本只保留最近 2 版,更早版本仅留索引与哈希(存档时间线的证据链不会因此断裂)。
纠错与举报(发现条目失效、署名有误或涉及侵权?)
提交举报 / 纠错
侵权举报经核验成立后,我们会即时下线该条目并删除已存的内容副本。