1. X
  2. Software Engineering Papers
Log inSign up
Software Engineering Papers
54.1K posts
Image
user avatar
Software Engineering Papers
@ComputerPapers
New Software Engineering papers from arxiv.org: design tools, software metrics. Thank you to arXiv for use of its open access interoperability.
Worldwide
arxiv.org/list/cs.SE/new
Joined April 2010
4
Following
2,529
Followers
RepliesRepliesMediaMedia
  • user avatar
    Software Engineering Papers
    @ComputerPapers
    25m
    IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu arxiv.org/abs/2608.02641 [饾殞饾殰.饾殏饾櫞 饾殞饾殰.饾櫚饾櫢]
    Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost. We present IR2Solve, an intermediate-representation-first autoformulation pipeline that uses a single semantic LLM call to produce a schema-constrained ModelIR, followed by two deterministic stages: verification and IR-to-solver compilation. ModelIR explicitly represents sets, parameters, variables, objectives, and constraints using restricted Python-like expression strings. A concrete scalar-constraint convention represents finite per-index constraint families as individual entries, reducing free-index and implicit-quantification errors while simplifying downstream verification and compilation. Across six cleaned optimization benchmarks, IR2
  • user avatar
    Software Engineering Papers
    @ComputerPapers
    1h
    Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation Atul Anand, Sourav Chattaraj arxiv.org/abs/2608.02639 [饾殞饾殰.饾殏饾櫞 饾殞饾殰.饾櫚饾櫢]
    Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how instruction-following degrades as such constraints accumulate. We introduce a benchmark that stacks 24 verifier-checked instructions, one to twenty at a time, and evaluate three production-tier LLMs (Claude Sonnet 4.6, GPT-5-mini, Gemini 2.5 Flash). Instruction-following degrades non-linearly: the follow rate falls from 96% to as low as 20%, driven by a structured and reproducible set of pairwise conflicts. A single "output JSON" constraint, for example, is jointly unsatisfiable with nine others. We then evaluate a training-free remedy: an instruction compiler that rewrites the stacked prompt in a single LLM call and is reused across queries. Its benefit is capability-graded. It recovers up to +11 points of follow rate for weaker models, which are also the models most often deployed at scale, while leaving strong
  • user avatar
    Software Engineering Papers
    @ComputerPapers
    2h
    Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems Beena arxiv.org/abs/2608.02638 [饾殞饾殰.饾殏饾櫞 饾殞饾殰.饾櫚饾櫢]
    Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI-CPS are used in several domains, including autonomous vehicles, industry, home automation, robotics, and healthcare. Being composed of hardware, AI components, and conventional modules, AI-CPS can exhibit technical debt (TD) that is peculiar and potentially more challenging than that of conventional systems. This thesis aims to characterize AI-CPS TD and propose approaches for its identification and repair. In a first phase, we characterize AI-CPS TD by analyzing AI ecosystems and AI-CPS repositories, as well as interviewing developers. Based on the acquired knowledge, we define approaches to identify and mitigate such TD. Finally, we plan to develop and validate an automated tool that supports agentic AI solutions to monitor, govern, and repay AI-CPS TD.
  • user avatar
    Software Engineering Papers
    @ComputerPapers
    2h
    Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song arxiv.org/abs/2608.02636 [饾殞饾殰.饾殏饾櫞 饾殞饾殰.饾櫚饾櫢]
    Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model. Yet it remains unclear when further evolution helps, how successful and failed trajectories shape revision, and whether extra test-time computation can recover the same gains. To address these questions, we present a controlled evaluation framework across five benchmarks and three models. Our primary study contains 42 feedback runs across 14 supported model-benchmark settings. Within each setting, we hold the executor and optimizer configuration, revision procedure, validation rule, and round budget fixed, while varying only the feedback shown to the optimizer: successes and failures (Normal), failures only, or successes only. Evolution is sparse: only 55 of 388 candidates establish byte-distinct validation bests. Validation-based selection chooses an evolved skill in 11 of 14 settings, nine of which improve released-test performance.
  • user avatar
    Software Engineering Papers
    @ComputerPapers
    14h
    ACEM: A Cost Estimation Model for Agentic Software Engineering Mohammad El-Ramly arxiv.org/abs/2608.02582 [饾殞饾殰.饾殏饾櫞]
    Traditional software cost estimation models, such as COCOMO II, Function Points, and Story Points, assume that development effort is primarily driven by human labor in design, coding, and testing. Agentic software engineering, where autonomous AI agents perform substantial implementation work and humans focus on planning, specification, and validation, challenges this assumption. New cost dimensions arise: large language model (LLM) token consumption across agent actions, Human-in-the-Loop (HITL) oversight effort, and infrastructure costs for agent orchestration and tooling. These costs are nondeterministic: identical tasks may consume different tokens, follow divergent reasoning paths, and require varying human correction, phenomena absent in traditional development. A new framework is needed to bridge standard sizing metrics with this cost structure. This paper proposes ACEM (Agentic Cost Estimation Model), which decomposes total agentic development cost into three additive dimension

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email
Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.
Advertisement
Advertisement