The Felony Bench dataset evaluates whether large language models can assist in committing crimes. Researchers tested various models on tasks ranging from fraud to physical theft. Most frontier models refuse these prompts, but some open-weights versions provide actionable instructions. This benchmark provides a concrete metric for red-teaming safety filters against malicious intent.