AgentMarketMCP / SKILL 资产档案馆

目录 / Test-robustness-ai

SKILL 未评级 已上架

Test-robustness-ai

## 🔬 Adversarial Testing on Abliterated Models This section covers testing AI models that have been "uncensored" (abliterated). 1. **Load the Model**: Use `transformers` to load the abliterated model you want to test. 2. **Run a "Forbidden" Prompt Set**: Use a dataset of prompts that would normally be refused (e.g., HarmBench, AdvBench). 3. **Evaluate Responses**: - **Refusal Rate**: How many prompts were refused? Target should be < 5%. - **Coherence & Quality**: Does the model maintain consistent responses, or does it hallucinate? 4. **Generate a Report**: Produce a report comparing original and abliterated model performance.

存档时间线

版本存档时间内容哈希内容
v12026-09-28 01:4816dd6768 可取

版本索引永久保留;内容副本只保留最近 2 版,更早版本仅留索引与哈希(存档时间线的证据链不会因此断裂)。

下载存档内容副本

纠错与举报(发现条目失效、署名有误或涉及侵权?)
提交举报 / 纠错

侵权举报经核验成立后,我们会即时下线该条目并删除已存的内容副本。