Dagstuhl Seminar 27501
Combining Learning and Reasoning for Programming Intelligence
( Dec 12 – Dec 17, 2027 )
Permalink
Organizers
- José Cambronero (Google - Austin, US)
- Isil Dillig (University of Texas - Austin, US)
- Devdatt Dubhashi (Chalmers University of Technology - Göteborg, SE)
- Vijay Ganesh (Georgia Institute of Technology - Atlanta, US)
- Adish Singla (MPI-SWS - Saarbrücken, DE)
Contact
- Marsha Kleinbauer (for scientific matters)
- Simone Schilke (for administrative matters)
Recently, large language models (LLMs) have demonstrated substantial progress across tasks in formal domains, including program synthesis, software verification, program analysis, testing, repair, and proof generation. Despite being trained with general next-token prediction, these models can now successfully generate code in general-purpose languages as well as producing proofs using both automated and interactive theorem provers [1, 2, 3].
While these generative models are highly capable, they lack the formal correctness guarantees required for critical software engineering and mathematical applications. To address these limitations, researchers are increasingly incorporating formal and symbolic reasoning into how these models are trained and executed [14]. For example, recent work explores reformulating tasks declaratively and using logical solvers to answer them [3], using compilers [4] or verifiers as automated feedback sources to scale test-time compute [5]. These advances are being explored concurrently, but often independently, across three disjoint research communities: AI for code, AI for math, and formal methods, while drawing expertise from both software engineering and natural language processing. Because these groups often work in isolation, there is a strong need to share successful and failed experimental methodologies, identify shared challenges, and build collaborations across these disciplines.
Focus of the Seminar
This Dagstuhl Seminar will focus on bridging these disjoint communities by structuring discussions around four primary technical areas that span the lifecycle of model training, inference, and execution.
First, we will examine reasoning for training, which includes curating program and proof datasets, and using symbolic tools as verifiers or reward models for reinforcement learning [6]. Second, we will address reasoning at test-time, exploring how to augment model prompts with formal logic context [7], and how to scale test-time compute using search and feedback from reasoning tools [13]. Third, we will study the formalization of inputs and outputs, investigating how to blend informal natural language with formal specifications to ensure the reliability and trustworthiness of LLM generation [8, 9]. Finally, we will cover the orchestration of symbolic and LLM-based tooling, analyzing agentic workflows [4, 10] and evaluating how to make use of generative models within these broader symbolic harnesses.
Goals and Expected Artifacts
The expected outcomes of the seminar are aimed at establishing common ground and aligning research agendas on critical needs like robust benchmarking [11, 12]. Participants will collaborate on two concrete deliverables: a co-authored vision paper outlining a shared research agenda and a public, community-maintained "awesome-style" GitHub resource repository hosting tools, datasets, and benchmarks for the broader research community.
References
- [1] Chakraborty, Saikat, et al. "Towards neural synthesis for SMT-assisted proof-oriented programming." 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 2025.
- [2] Swamy, Nikhil, et al. "Secure distributed programming with value-dependent types." ACM SIGPLAN Notices 46.9 (2011): 266-278.
- [3] Ye, Xi, et al. "Satlm: Satisfiability-aided language models using declarative prompting." Advances in Neural Information Processing Systems 36 (2023): 45548-45580.
- [4] Rondon, Pat, et al. "Evaluating agent-based program repair at google." 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2025.
- [5] Snell, Charlie, et al. "Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning." International Conference on Learning Representations. Vol. 2025. 2025.
- [6] Guo, Daya, et al. "Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning." arXiv preprint arXiv:2501.12948 (2025).
- [7] Yang, Kaiyu, et al. "Leandojo: Theorem proving with retrieval-augmented language models." Advances in Neural Information Processing Systems 36 (2023): 21573-21612.
- [8] Le-Cong, Thanh, Bach Le, and Toby Murray. "Can LLMs reason about program semantics? a comprehensive evaluation of LLMs on formal specification inference." Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
- [9] Misu, Md Rakib Hossain, et al. "Towards ai-assisted synthesis of verified dafny methods." Proceedings of the ACM on Software Engineering 1.FSE (2024): 812-835.
- [10] Yang, John, et al. "Swe-agent: Agent-computer interfaces enable automated software engineering." Advances in Neural Information Processing Systems 37 (2024): 50528-50652.
- [11] Chandra, Satish. "Benchmarks for AI in Software Engineering." BLOG@CACM, 2025, cacm.acm.org/blogcacm/benchmarks-for-ai-in-software-engineering/
- [12] Zhu, Yuxuan, et al. "Establishing best practices in building rigorous agentic benchmarks." Advances in Neural Information Processing Systems 38 (2026).
- [13] Jana, Prithwish, et al. "Proofbridge: Auto-formalization of natural language proofs in lean via joint embeddings." International Conference on Learning Representations. Vol. 2026. 2026.
- [14] Yang, Kaiyu, et al. "Formal reasoning meets llms: Toward ai for mathematics and verification." Communications of the ACM 69.3 (2026): 66-73.
José Cambronero, Isil Dillig, Devdatt Dubhashi, Vijay Ganesh, and Adish Singla
Classification
- Artificial Intelligence
- Programming Languages
Keywords
- Neurosymbolic AI
- Artificial Intelligence
- AI for Math
- AI for Code
- Formal Methods

Creative Commons BY 4.0
