10.57647/ijm2c.2027.1702.12

SpecPatch: Specification-Aware Fault Localization and Constraint-Guided Program Repair for C/C++ Systems

  1. Department of Computer Engineering and Information Technology, Qa.C., Islamic Azad University, Qazvin, Iran
  2. Department of Computer Engineering, Iran University of Science and Technology, Tehran, Iran
  3. Department of Computer Science, Faculty of Mathematical Sciences, University of Tabriz, Tabriz, Iran

Received: 04-07-2026

Revised: 11-07-2026

Accepted: 12-07-2026

Published Online: 15-07-2026

How to Cite

Ghanati, A., Parsa, S., & Izadkhah, H. (2025). SpecPatch: Specification-Aware Fault Localization and Constraint-Guided Program Repair for C/C++ Systems. International Journal of Mathematical Modelling & Computations. https://doi.org/10.57647/ijm2c.2027.1702.12

Abstract

Automated program repair (APR) for C/C++ remains challenging because fault-localization signals are uncertain, candidate spaces are large, validation is costly, and learned, neural, or large language model (LLM)-based repair decisions are often difficult to audit. This paper presents SpecPatch, an uncertainty-aware fuzzy learning-to-rank framework that prioritizes explicit repair-pattern families before candidate generation. Unlike open-ended neural or LLM-based code generation, SpecPatch does not synthesize arbitrary source code. Instead, it represents each suspicious code region using 40 Boolean fault-local and contextual features and ranks 13 auditable abstract syntax tree (AST)-level repair-pattern families, including nine adapted families and four proposed C/C++ patterns: Library-Function Repair, Enum Repair, For-Statement Repair, and Neighbourhood Repair. The central contribution is a fuzzy inference layer that converts overlapping syntactic, data-dependence, control-context, API-call, enumeration, loop-header, and structural-risk evidence into interpretable repair-suitability scores. These scores are fused with pairwise learning-to-rank relevance, while syntactic, type, scope, and structural preconditions remain hard validity constraints. Evaluation on 2,244 held-out pattern-selection samples and a 69-defect C/C++ benchmark shows that the balanced fuzzy-LTR configuration improves rank-aware ordering and reduces search effort. Compared with the LTR-only SpecPatch baseline, it improves Recall@1 from 0.45 to 0.50, Recall@3 from 0.74 to 0.80, and NDCG@5 from 0.78 to 0.84. It also reduces generated candidates from 914 to 620 per defect, compilation attempts from 337 to 238, and mean runtime from 49.38 to 41.70 minutes. Observed differences in plausible and correct-first repairs are reported as secondary outcomes and are not used as the main superiority claim. LLM-based repair is treated as a contextual reference point; the primary contribution of SpecPatch is a complementary, auditable, reproducible, and cost-aware repair-pattern ranking method rather than unrestricted code generation.

Keywords

  • Automated program repair,
  • C/C++,
  • Fuzzy logic,
  • Learning to rank,
  • Interpretable machine learning,
  • Repair patterns,
  • Patch ranking,
  • Program analysis

References

  1. Gazzola L, Micucci D, Mariani L. Automatic software repair: A survey. IEEE Transactions on Software Engineering. 2019;45(1):34–67.
  2. Le Goues C, Nguyen T, Forrest S, Weimer W. GenProg: A generic method for automatic software repair. IEEE Transactions on Software Engineering. 2012;38(1):54–72. doi: https://doi.org/10.1109/TSE.2011.104
  3. Qi Z, Long F, Achour S, Rinard M. An analysis of patch plausibility and correctness for generate-and-validate patch generation systems. In: Proceedings of the 2015 International Symposium on Software Testing and Analysis (ISSTA); 2015. p. 24–36. doi: https://doi.org/10.1145/2771783.2771791
  4. Long F, Rinard M. Staged program repair with condition synthesis. In: Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE); 2015. p. 166–178. doi: https://doi.org/10.1145/2786805.2786811
  5. Long F, Rinard M. Automatic patch generation by learning correct code. In: Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI); 2016. p. 298–312.
  6. Chen Z, Kommrusch S, Tufano M, Pouchet L, Poshyvanyk D, Monperrus M. SequenceR: Sequence-to-sequence learning for end-to-end program repair. IEEE Transactions on Software Engineering. 2021;47(9):1940–1956.
  7. Lutellier T, Pham H, Pang L, Li Y, Wei M, Tan L. CoCoNuT: Combining context-aware neural translation models using an ensemble for program repair. In: Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA); 2020. p. 101–114.
  8. Feng Z, Guo D, Tang D, Duan N, Feng X, Gong M, et al. CodeBERT: A pre-trained model for programming and natural languages. In: Findings of the Association for Computational Linguistics: EMNLP 2020; 2020. p. 1536–1547.
  9. Xia CS, Zhang L. Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. arXiv:2304.00385. 2023.
  10. Yin X, Ni C, Wang S, Li Z, Zeng L, Yang X. ThinkRepair: Self-directed automated program repair. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA); 2024. doi: https://doi.org/10.1145/3650212.3680359
  11. Bouzenia I, Devanbu P, Pradel M. RepairAgent: An autonomous, LLM-based agent for program repair. In: Proceedings of the 47th IEEE/ACM International Conference on Software Engineering (ICSE); 2025.
  12. Majd A, Vahidi-Asl M, Khalilian A, Baraani-Dastjerdi A, Zamani B. Code4Bench: A multidimensional benchmark of Codeforces data for different program analysis techniques. Journal of Computer Languages. 2019;53:38–52. doi: https://doi.org/10.1016/j.cola.2019.03.006
  13. Qi Y, Mao X, Lei Y, Dai Z, Wang C. The strength of random search on automated program repair. In: Proceedings of the 36th International Conference on Software Engineering (ICSE); 2014. p. 254–265. doi: https://doi.org/10.1145/2568225.2568254
  14. Kim D, Nam J, Song J, Kim S. Automatic patch generation learned from human-written patches. In: Proceedings of the 35th International Conference on Software Engineering (ICSE); 2013. p. 802–811.
  15. Xuan J, Martinez M, DeMarco F, Clément M, Lamelas Marcote S, Durieux T, et al. Nopol: Automatic repair of conditional statement bugs in Java programs. IEEE Transactions on Software Engineering. 2017;43(1):34–55.
  16. Liu K, Koyuncu A, Kim D, Bissyandé TF. TBar: Revisiting template-based automated program repair. In: Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA); 2019. p. 31–42.
  17. White M, Tufano M, Martinez M, Monperrus M, Poshyvanyk D. Sorting and transforming program repair ingredients via deep learning code similarities. In: Proceedings of the 26th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER); 2019. p. 479–490.
  18. Silva A, Fang S, Monperrus M. RepairLLaMA: Efficient representations and fine-tuned adapters for program repair. arXiv:2312.15698. 2023.
  19. Jiang N, Lutellier T, Tan L. CURE: Code-aware neural machine translation for automatic program repair. In: Proceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE); 2021. p. 1161–1173.
  20. Xia CS, Zhang L. Less training, more repairing please: Revisiting automated program repair via zero-shot learning. In: Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE); 2022.
  21. Sun J, Li F, Qi X, Zhang H, Jiang J. Empirical evaluation of large language models in automated program repair. arXiv:2506.13186. 2025.
  22. Majd A, Vahidi-Asl M, Khalilian A, Baraani-Dastjerdi A, Zamani B. Code4Bench: A multidimensional benchmark of Codeforces data for different program analysis techniques. Zenodo. ver. 1.0.0. 2019 Mar. doi: https://doi.org/10.5281/zenodo.2582968
  23. Code4Bench Project. Code4Bench companion repository: schema, field definitions, and extraction scripts. GitHub repository. Accessed July 2026.
  24. Burges C, Shaked T, Renshaw E, Lazier A, Deeds M, Hamilton N, et al. Learning to rank using gradient descent. In: Proceedings of the 22nd International Conference on Machine Learning (ICML); 2005. p. 89–96. doi: https://doi.org/10.1145/1102351.1102363
  25. Burges CJC. From RankNet to LambdaRank to LambdaMART: An overview. Microsoft Research Technical Report MSR-TR-2010-82. 2010.
  26. Järvelin K, Kekäläinen J. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems. 2002;20(4):422–446. doi: https://doi.org/10.1145/582415.582418
  27. Abreu R, Zoeteweij P, Golsteijn R, van Gemund AJC. A practical evaluation of spectrum-based fault localization. Journal of Systems and Software. 2009;82(11):1780–1792. doi: https://doi.org/10.1016/j.jss.2009.06.035
  28. Jones JA, Harrold MJ. Empirical evaluation of the Tarantula automatic fault-localization technique. In: Proceedings of the 20th IEEE/ACM International Conference on Automated Software Engineering (ASE); 2005. p. 273–282. doi: https://doi.org/10.1145/1101908.1101949
  29. Pearson S, Campos J, Just R, Fraser G, Abreu R, Ernst MD, et al. Evaluating and improving fault localization. In: Proceedings of the 39th International Conference on Software Engineering (ICSE); 2017. p. 609–620. doi: https://doi.org/10.1109/ICSE.2017.62
  30. Hui B, Yang J, Cui Z, Yang J, Liu D, Zhang L, et al. Qwen2.5-Coder Technical Report. arXiv:2409.12186. 2024.
  31. Qwen Team. Qwen2.5-Coder-7B-Instruct. Hugging Face model card. Accessed July 2026.
  32. Zadeh LA. Fuzzy sets. Information and Control. 1965;8(3):338–353.
  33. Mamdani EH, Assilian S. An experiment in linguistic synthesis with a fuzzy logic controller. International Journal of Man-Machine Studies. 1975;7(1):1–13.
  34. Takagi T, Sugeno M. Fuzzy identification of systems and its applications to modeling and control. IEEE Transactions on Systems, Man, and Cybernetics. 1985;SMC-15(1):116–132.
  35. Jang JSR. ANFIS: Adaptive-network-based fuzzy inference system. IEEE Transactions on Systems, Man, and Cybernetics. 1993;23(3):665–685.
  36. Zhu Q, Sun Z, Xiao Y, Zhang W, Yuan K, Xiong Y, et al. A syntax-guided edit decoder for neural program repair. In: Proceedings of the 29th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE); 2021. p. 341–353. doi: https://doi.org/10.1145/3468264.3468544
  37. Xia CS, Wei Y, Zhang L. Automated program repair in the era of large pre-trained language models. In: Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE); 2023. doi: https://doi.org/10.1109/ICSE48619.2023.00129
  38. Zhang Q, Fang C, Zhang T, Yu B, Sun W, Chen Z. GAMMA: Revisiting template-based automated program repair via mask prediction. In: Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE); 2023. p. 535–547. doi: https://doi.org/10.1109/ASE56229.2023.00063
  39. Xia CS, Ding Y, Zhang L. The plastic surgery hypothesis in the era of large language models. In: Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE); 2023. doi: https://doi.org/10.1109/ASE56229.2023.00047
  40. Huang K, Zhang J, Meng X, Liu Y. Template-guided program repair in the era of large language models. In: Proceedings of the 47th IEEE/ACM International Conference on Software Engineering (ICSE); 2025. doi: https://doi.org/10.1109/ICSE55347.2025.00030
  41. Chen Y, Wu J, Ling X, Li C, Rui Z, Luo T, et al. When large language models confront repository-level automatic program repair: How well do they perform? In: Proceedings of the 46th IEEE/ACM International Conference on Software Engineering Companion (ICSE Companion); 2024. doi: https://doi.org/10.1145/3639478.3647633
  42. Yang B, Cai Z, Liu F, Le B, Zhang L, Bissyandé TF, et al. A survey of LLM-based automated program repair: Taxonomies, design paradigms, and applications. arXiv:2506.23749. 2025.