Research
Browse Papers by Research Topic
Explore publications organized by research themes — from LLM benchmarking and formal verification to AI testing and software security.
Showing 42 papers
ICML '26
2026
ICML '26
2026
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Benchmark
AAAI '25
2025
DomainEval: An Auto-Constructed Benchmark for Multi-Domain Code Generation
Benchmark
ACL '25
2025
CruxEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
Benchmark
ASE '24
2024
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
Benchmark
ASE '24
2024
Internetware '25
2025
arXiv
2025
arXiv
2025
arXiv
2025
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
Benchmark
ACL '25
2025
From Informal to Formal — Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
10k+ HuggingFace downloads · 37k+ social media reads
Formal Methods
CAV '24
2024
FM '26
2026
arXiv
2026
arXiv
2026
LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation
Formal Methods
ICSE '26
2026
TOSEM '26
2026
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
LLM for SE
ACL '26
2026
Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation
LLM for SE
FSE '26
2026
TOSEM '26
2026
When Retrieval Augmentation Meets API Documentation: Can LLMs Code with Less-Common Libraries?
LLM for SE
ASEJ '25
2025
ASE '26
2026
Doc2Feat-bench: Evaluating Documentation-Driven Feature Addition
LLM for SE
arXiv
2025
arXiv
2025
arXiv
2025
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
LLM for SE
ICSE '22
2022
TOSEM '22
2022
ASE '24
2024
MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing
SE for AI
FSE '25
2025
TOSEM '23
2023
TOSEM '25
2025
AAAI '25
2025
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
SE for AI
USENIX Sec '22
2022
RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix
Security
USENIX Sec '21
2021
ReDoSHunter: A Combined Static and Dynamic Approach for Regular Expression DoS Detection
Security
TASE '24
2024
Fuzzing for Stateful Protocol Implementations: Are We There Yet?
Security
ASE '25
2025
Vulnerability-Affected Versions Identification: How Far Are We?
Security
APSEC '24
2024
SDEFL: A Lightweight Fault Detection and Localization Method for Deep Neural Networks
Security
FSE '23
2023
Understanding the Bug Characteristics and Fix Strategies of Federated Learning Systems
LLM for SE
ICSE '21
2021
ICCD '19
2019