Three large language models screened abstracts and full-text articles across five diverse systematic review projects. This research tests whether AI can replace the manual labor of identifying eligible clinical trials. The results provide a benchmark for using LLMs to accelerate evidence-based medicine. Practitioners can now quantify the safety of automating study selection.