# Hugging Face — CyberSecEval 2 - A Comprehensive Evaluation Framework for Cybersecurity Risks and Capabilities of Large Language Models

- Company: Hugging Face (huggingface.co)
- Announced: 2024-05-24T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Source: https://huggingface.co/blog/leaderboard-llamaguard
- Record: https://forck.live/items/1883-cyberseceval-2-a-comprehensive-evaluation-framework-for-cybersecurity-risks

The blog post introduces CyberSecEval 2, a comprehensive evaluation framework for cybersecurity risks and capabilities of Large Language Models (LLMs). It includes benchmarks for insecure code generation, prompt injection, compliance with cyber attack requests, code interpreter abuse, and automated offensive cybersecurity capabilities. The post also provides key insights from evaluating state-of-the-art LLMs, noting a decrease in compliance with cyber attack requests since the first version, but highlighting ongoing challenges with prompt injection and interpreter abuse.

## Evidence

Verbatim from https://huggingface.co/blog/leaderboard-llamaguard:

> CyberSecEval 2, which assesses an LLM's susceptibility to code interpreter abuse, offensive cybersecurity capabilities, and prompt injection attacks, comes into play to provide a more comprehensive evaluation of LLM cybersecurity risks.

---

Record: https://forck.live/items/1883-cyberseceval-2-a-comprehensive-evaluation-framework-for-cybersecurity-risks
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
