# PatchEval **Repository Path**: ByteDance/PatchEval ## Basic Information - **Project Name**: PatchEval - **Description**: No description available - **Primary Language**: Unknown - **License**: Apache-2.0 - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2025-11-19 - **Last Updated**: 2026-07-25 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README
Logo

A New Benchmark for Evaluating LLMs / Agents on Patching Real-World Vulnerabilities

Python License

--- ## ๐Ÿ“ข News * **[2026/07/24]** We release **PatchEval-Verified**, with updated Docker evaluation environments for 230 CVEs. In the original PatchEval, some PoC tests were adapted from project regression tests and were tied too closely to particular patch implementations. We revised these tests to evaluate whether a vulnerability is fixed, rather than requiring a specific fix, improving robustness and reducing false negatives. * **[2025/11/18]** PatchEval is released as a benchmark for evaluating Large Language Models and agents on real-world vulnerability repair. ## ๐Ÿ‘‹ Overview **PatchEval-Verified** evaluates LLMs and coding agents on automated repair of real-world vulnerabilities. This repository provides **230 CVE cases with Docker-based evaluation environments**, covering vulnerabilities reported between 2015 and 2025 across Go, JavaScript, and Python. Every image contains the vulnerable repository and a validation entrypoint. The main improvement over the original PatchEval release is the upgraded evaluation harness: * **Less implementation coupling:** a valid repair should not need to reproduce the official patch or one specific implementation strategy. * **Stronger vulnerability validation:** the validation tests are designed to check whether the vulnerability remains exploitable after applying the generated patch. * **More robust scoring:** strengthened test logic reduces accidental passes and rejection of semantically correct alternative fixes. The repository also provides a streamlined two-stage workflow: ```text 1. Generate patches with a CLI coding agent patcheval/exp_agent/run_infer.sh 2. Evaluate generated patches in the verified Docker environments patcheval/exp_agent/run_eval.sh -> patcheval/evaluation/run_evaluation.py ``` The included agent adapters are: ```text Codex CLI OpenCode TraeCLI ``` The figure below illustrates the overall design of the original PatchEval benchmark. This repository focuses on the 230 Docker-verified cases and the dynamic evaluation workflow.

## ๐Ÿ’ป Getting Started ### Requirements * **Operating System**: Linux. * **Python**: Python 3.10+; Python 3.12 is recommended. * **Docker**: Docker must be installed and available to the current user. * **CPU**: 16 or more cores are recommended for parallel experiments. * **Disk Storage**: Reserve at least 500 GB if pulling all 230 images locally. ### Setup ```bash git clone https://github.com/bytedance/PatchEval.git cd PatchEval conda create -n patcheval python=3.12 conda activate patcheval pip install -r requirements.txt ``` Verify both Docker and the Python Docker SDK: ```bash docker version python -c "import docker; print(docker.__version__)" ``` ## ๐Ÿ“œ Repository Structure ```text ./ โ”œโ”€โ”€ docs/ โ”œโ”€โ”€ patcheval/ โ”‚ โ”œโ”€โ”€ datasets/ โ”‚ โ”‚ โ””โ”€โ”€ patcheval_verified.json # metadata and image URL for 230 verified cases โ”‚ โ”œโ”€โ”€ evaluation/ โ”‚ โ”‚ โ”œโ”€โ”€ README.md โ”‚ โ”‚ โ”œโ”€โ”€ run_evaluation.py # Docker-based patch evaluator โ”‚ โ”‚ โ”œโ”€โ”€ utils.py โ”‚ โ”‚ โ””โ”€โ”€ example_patch.json โ”‚ โ””โ”€โ”€ exp_agent/ โ”‚ โ”œโ”€โ”€ agents/ # Codex, OpenCode, and TraeCLI adapters โ”‚ โ”œโ”€โ”€ patch_agent_runner.py # patch-generation-only runner โ”‚ โ”œโ”€โ”€ process_data.py # converts generated patches to evaluation JSONL โ”‚ โ”œโ”€โ”€ run_infer.sh # patch generation entrypoint โ”‚ โ”œโ”€โ”€ run_eval.sh # evaluation entrypoint โ”‚ โ””โ”€โ”€ README.md โ”œโ”€โ”€ scripts/ โ”‚ โ”œโ”€โ”€ download_images.py โ”‚ โ””โ”€โ”€ images.txt # list of the 230 Docker images โ”œโ”€โ”€ README.md โ””โ”€โ”€ requirements.txt ``` ## ๐Ÿ“Š Verified Dataset and Docker Images ### Dataset The metadata used by both patch generation and evaluation is: ```text patcheval/datasets/patcheval_verified.json ``` It contains 230 entries with fields including: ```text cve_id cve_description cwe_info repo patch_url programing_language vul_func fix_func image_url ``` > [!NOTE] > `programing_language` is the field name used by the released dataset and is intentionally preserved for compatibility. The agent runner uses `cve_description`, `repo`, and `image_url` to construct the task and runtime environment. Reference information such as `patch_url` and `fix_func` is not included in the agent prompt and should not be used as a repair source. ### Docker Images `scripts/images.txt` contains the same 230 image references recorded by `image_url` in the verified dataset. Image names follow this form: ```text ghcr.io/patcheval-cve/patcheval-cve:cve-- ``` Pull all images with: ```bash cd scripts python download_images.py ``` The script pulls images concurrently. Make sure the images required by your selected cases are available locally before evaluation. ## ๐Ÿš€ CLI Agent Patch Generation and Evaluation The verified workflow has been smoke-tested with the included `codex`, `opencode`, and `traecli` adapters. All adapters use the same two-step interface: ```bash conda activate patcheval cd patcheval/exp_agent bash run_infer.sh # generate patches bash run_eval.sh # evaluate generated patches ``` Start with one case before scaling to the complete dataset. ### Codex smoke test ```bash conda activate patcheval cd patcheval/exp_agent export CODEX_BIN=/path/to/codex export CODEX_CONFIG=/path/to/codex-home/.config.toml LIMIT=1 CONCURRENCY=1 bash run_infer.sh codex codex_smoke MAX_WORKERS=1 bash run_eval.sh codex_smoke ``` ### Full 230-case run ```bash conda activate patcheval cd patcheval/exp_agent CONCURRENCY=10 bash run_infer.sh codex codex_full MAX_WORKERS=4 bash run_eval.sh codex_full ``` `CONCURRENCY` controls parallel patch-generation containers, while `MAX_WORKERS` controls parallel evaluation containers. Replace `codex` with `opencode` or `traecli` to use another adapter. See [patcheval/exp_agent/README.md](patcheval/exp_agent/README.md) for the required agent-specific environment variables. ## ๐Ÿงช Standalone Patch Evaluation To evaluate patches generated by another system, prepare a JSON or JSONL file whose entries contain: ```json { "cve": "CVE-2021-23376", "fix_patch": "diff --git ..." } ``` Then run: ```bash cd patcheval/evaluation python run_evaluation.py \ --output my_run \ --patch_file /path/to/patches.jsonl \ --input_file ../datasets/patcheval_verified.json \ --max_workers 4 \ --log_level INFO ``` For every patch, the evaluator: 1. starts the Docker image specified by the case's `image_url`; 2. mounts the submitted patch at `/workspace/fix.patch`; 3. runs `cd /workspace && bash fix-run.sh` with a 600-second command timeout; 4. classifies failures heuristically as patch-application, compilation, validation, or execution errors; 5. removes the temporary container. A repair passes only when `fix-run.sh` exits successfully. See [patcheval/evaluation/README.md](patcheval/evaluation/README.md) for details. `run_infer.sh` returns a non-zero exit code if any selected case fails, while preserving all generated patches and logs. Failed generation cases are represented by empty patch files and remain in the evaluation input, where they are counted as failed repairs. ## ๐Ÿ“ Outputs A patch-generation run creates: ```text patcheval/exp_agent/agent_runs/-/ โ”œโ”€โ”€ patches/ # one patch file per CVE โ”œโ”€โ”€ .work/ # prompts and agent stdout/stderr โ”œโ”€โ”€ results.jsonl # per-case generation status โ”œโ”€โ”€ run_metadata.json โ””โ”€โ”€ summary.json # generated/failed counts ``` `run_eval.sh` converts these patches to: ```text patcheval/exp_agent/eval_inputs/.jsonl ``` Evaluation results are written to: ```text patcheval/evaluation/evaluation_output/results// โ”œโ”€โ”€ run_evaluation.log โ”œโ”€โ”€ summary_report.txt โ”œโ”€โ”€ summary.json โ””โ”€โ”€ logs// โ”œโ”€โ”€ fix.patch โ””โ”€โ”€ success_output.log or error_output.log ``` ## ๐Ÿ“ˆ Results on PatchEval-Verified Coming soon... ## ๐Ÿš€ Contributions We welcome issues, bug reports, improvements to the verified validation environments, agent adapters, and reproducibility scripts. ## ๐Ÿ“– Citation If you find PatchEval useful for your research and applications, please cite: ```bibtex @misc{wei2025patcheval, title={PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities}, author={Zichao Wei and Jun Zeng and Ming Wen and Zeliang Yu and Kai Cheng and Yiding Zhu and Jingyi Guo and Shiqi Zhou and Le Yin and Xiaodong Su and Zhechao Ma}, year={2025}, eprint={2511.11019}, archivePrefix={arXiv}, primaryClass={cs.CR}, url={https://arxiv.org/abs/2511.11019}, } ``` ## โœ๏ธ License This project is licensed under the Apache License 2.0. See the [LICENSE](LICENSE) file for details.