{
  "manuscript": {
    "version": "advisor-v3.2",
    "filename": "TUCO-JISE-advisor-v3.2.pdf",
    "sha256": "c716b17c233d6df968a793c8f54a8f06046939827d4ada4c19a111ef5aaf1525",
    "pages": 18,
    "references": 33,
    "citation_groups": 49,
    "citation_links": 54,
    "date": "2026-09-23",
    "source_zip_sha256": "0ebb16d3d7c088c6924884053247050a8db07dee93be0b9164f3da15cc09771c"
  },
  "references": [
    {
      "number": 1,
      "key": "zhou2024webarena",
      "bib": {
        "author": "Zhou, Shuyan and Xu, Frank F. and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham",
        "title": "WebArena: A Realistic Web Environment for Building Autonomous Agents",
        "booktitle": "Proceedings of the International Conference on Learning Representations (ICLR)",
        "year": "2024",
        "url": "https://openreview.net/forum?id=oKn9c6ytLx"
      },
      "url": "https://openreview.net/forum?id=oKn9c6ytLx",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/zhou-2024-webarena/auto/zhou-2024-webarena_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "bf59c728c47da388d71bfe2c317c5882368e6f73ad18b0dadefa8bcb262e2693",
      "pages": 22,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916155041",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "resolved",
        "note": "原文直接支持自架、開源、自然語言驅動的瀏覽器環境。v3.1 寫「made … the standard subject」，單篇原文不能證明整個領域的標準地位；v3.2 已改為「introduced a self-hostable environment built from open-source applications for browser agents driven by natural-language instructions」，與原文 §1、§2 的環境設計敘述一致。",
        "passages": [
          {
            "page": 2,
            "blocks": [
              19
            ],
            "label": "Introduction：四個自架應用與 Docker",
            "quote": null
          },
          {
            "page": 3,
            "blocks": [
              13
            ],
            "label": "§2：以開源軟體建立環境",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2307.13854v4",
      "fulltext_url": "https://arxiv.org/pdf/2307.13854v4",
      "record_url": "https://arxiv.org/abs/2307.13854v4",
      "evidence": [
        {
          "id": "R01-E01",
          "number": 1,
          "page": 2,
          "label": "Introduction：四個自架應用與 Docker",
          "image": "assets/ref-01-e01.png",
          "source_sha256": "bf59c728c47da388d71bfe2c317c5882368e6f73ad18b0dadefa8bcb262e2693",
          "image_sha256": "b81631d541237bf8846168d932e6cc8f288b1525a715f216cca0659420047f5a",
          "text": "We introduce WebArena, a realistic and reproducible web environment designed to facilitate the\ndevelopment of autonomous agents capable of executing tasks (§2). An overview of WebArena\nis in Figure 1. Our environment comprises four fully operational, self-hosted web applications,\neach representing a distinct domain prevalent on the internet: online shopping, discussion forums,\ncollaborative development, and business content management. Furthermore, WebArena incorporates\nseveral utility tools, such as map, calculator, and scratchpad, to best support possible human-like task\nexecutions. Lastly, WebArena is complemented by an extensive collection of documentation and\nknowledge bases that vary from general resources like English Wikipedia to more domain-specific\nreferences, such as manuals for using the integrated development tool (Fan et al., 2022). The content\npopulating these websites is extracted from their real-world counterparts, preserving the authenticity\nof the content served on each platform. We deliver the hosting services using Docker containers with\ngym-APIs (Brockman et al., 2016), ensuring both the usability and the reproducibility of WebArena.",
          "bbox": [
            107.03199768066406,
            323.4474182128906,
            506.24383544921875,
            455.44287109375
          ],
          "mode": "paragraph-crop",
          "width": 879,
          "height": 291
        },
        {
          "id": "R01-E02",
          "number": 1,
          "page": 3,
          "label": "§2：以開源軟體建立環境",
          "image": "assets/ref-01-e02.png",
          "source_sha256": "bf59c728c47da388d71bfe2c317c5882368e6f73ad18b0dadefa8bcb262e2693",
          "image_sha256": "cceea52dde988419fd76324757dfd00fb5604e514e833eb0b10bbadb720082a7",
          "text": "challenges such as bots being subject to CAPTCHAs, unpredictable content modifications, and\nconfiguration changes, which obstruct a fair comparison across different systems over time. We\nachieve realism by using open-source libraries that underlie many in-use sites from several popular\ncategories and importing data to our environment from their real-world counterparts.",
          "bbox": [
            107.5,
            276.6012878417969,
            504.6646728515625,
            320.51507568359375
          ],
          "mode": "paragraph-crop",
          "width": 875,
          "height": 98
        }
      ],
      "identity_image": "assets/ref-01-identity.png"
    },
    {
      "number": 2,
      "key": "teoh2026webtestpilot",
      "bib": {
        "author": "Teoh, Xiwen and Lin, Yun and Nguyen, Duc-Minh and Ren, Ruofei and Zhang, Wenjie and Dong, Jin Song",
        "title": "WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements",
        "journal": "Proceedings of the ACM on Software Engineering",
        "volume": "3",
        "number": "FSE",
        "pages": "1933--1956",
        "year": "2026",
        "doi": "10.1145/3797115"
      },
      "url": "https://doi.org/10.1145/3797115",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/webtestpilot_2602.11724/auto/webtestpilot_2602.11724_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "948830ec618057f1376ef2fb84a07f564d971b68ad1595002557ecc2e6699892",
      "pages": 24,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816111658",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "§4 說明符號化 GUI 與 DSL 斷言；§5.1.1 列出 BookStack、PrestaShop 等四個應用；§5.1.4 明確說每項需求設計一個人工缺陷。這些段落支撐本文的四處引用，並未證明 Tuco 與其使用相同版本或同一批缺陷。",
        "passages": [
          {
            "page": 4,
            "blocks": [
              2
            ],
            "label": "§4 方法概述：符號化與 DSL",
            "quote": null
          },
          {
            "page": 12,
            "blocks": [
              7
            ],
            "label": "§5.1.1：應用選取條件及 BookStack／PrestaShop",
            "quote": null
          },
          {
            "page": 13,
            "blocks": [
              14
            ],
            "label": "§5.1.4：每項需求一個人工缺陷",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2602.11724v3",
      "fulltext_url": "https://arxiv.org/pdf/2602.11724v3",
      "record_url": "https://arxiv.org/abs/2602.11724v3",
      "evidence": [
        {
          "id": "R02-E01",
          "number": 2,
          "page": 4,
          "label": "§4 方法概述：符號化與 DSL",
          "image": "assets/ref-02-e01.png",
          "source_sha256": "948830ec618057f1376ef2fb84a07f564d971b68ad1595002557ecc2e6699892",
          "image_sha256": "5fab9913c114b838188a2a12413a49ec0a7d5a90248cd9b10d59a18d8651ff05",
          "text": "Specifically, given a natural language test requirement, WebTestPilot decomposes it into 𝑛\n(condition, action, expectation) steps. For each step, WebTestPilot translates the condition and\nexpectation into pre- and post-condition assertions. It then applies symbolization to extract relevant\nUI components as symbols, which are composed via a DSL to construct executable assertions\nsatisfying the specified constraints. To support cross-state reasoning, WebTestPilot uses page\nreidentification to detect revisited pages and maintain a structured history of test states.",
          "bbox": [
            45.02899932861328,
            323.5299987792969,
            440.6706848144531,
            396.5093078613281
          ],
          "mode": "paragraph-crop",
          "width": 871,
          "height": 162
        },
        {
          "id": "R02-E02",
          "number": 2,
          "page": 12,
          "label": "§5.1.1：應用選取條件及 BookStack／PrestaShop",
          "image": "assets/ref-02-e02.png",
          "source_sha256": "948830ec618057f1376ef2fb84a07f564d971b68ad1595002557ecc2e6699892",
          "image_sha256": "1af1289c45f6b5e30bf7dd123ad01be6ebcae0e330a584411cef9fe1ff6cf566",
          "text": "5.1.1\nWeb Applications. We search GitHub for open-source web applications and select those based\non five criteria: (1) popularity, with ≥5,000 stars; (2) active development, with >50 contributors and\n>1,000 commits, and a commit in the past month; (3) maturity, publicly available for >5 years; (4)\npractical relevance, indicated by active deployment, recognizable domain or organization, commer-\ncial support, or adoption by well-known entities; and (5) user-facing documentation describing\ncore features. We select the following four web applications:\n• BookStack [8]: A hierarchical documentation management platform with rich text editing.\n• Indico [24]: An event manager for conferences, meetings, and lectures.\n• InvoiceNinja [25]: A business-oriented invoicing platform with multi-step workflows.\n• PrestaShop [50]: A full-stack e-commerce platform with store management feature.\nWe package the applications into reproducible Docker Compose environments.",
          "bbox": [
            45.327999114990234,
            420.598876953125,
            442.3524169921875,
            555.0232543945312
          ],
          "mode": "paragraph-crop",
          "width": 875,
          "height": 297
        },
        {
          "id": "R02-E03",
          "number": 2,
          "page": 13,
          "label": "§5.1.4：每項需求一個人工缺陷",
          "image": "assets/ref-02-e03.png",
          "source_sha256": "948830ec618057f1376ef2fb84a07f564d971b68ad1595002557ecc2e6699892",
          "image_sha256": "072a8488dcfd1ea6d4dd1afd43a0c87dfd6c2cc63f8cba9595a7f905edf0814a",
          "text": "5.1.4\nInjected Bugs. We design a single artificial bug bug𝑖𝑗: S →S for each test requirement\n𝐷𝑖𝑗. These bugs induce incorrect behaviors while ensuring stable and reproducible experiments by\nlocking application versions. To ensure realism, we examine closed GitHub issues labeled \"Bug\"\nfrom each application repository. From a total of 2,043 issues, we randomly sample 10%. We perform\nopen coding on the titles and descriptions of the sampled issues to identify meaningful labels, and\nthen conduct a thematic analysis to group these labels into broader bug categories. Two co-authors\nindependently perform the analysis, with a third resolving any disagreements. We exclude crash\nbugs and purely cosmetic bugs (e.g., layout or positioning issues) that do not affect functionality, as\nprior work has already addressed them. Based on our analysis, we focus on four categories:\n• Missing UI elements: Required interface components are absent, breaking feature functionality.\nFor example, in prestashop/#22170, the \"Configure\" button is missing for newly installed modules.\n• Data inconsistency: Information shown to the user does not match expected values. For example,\nin indico/#5197, the category search results include items that were previously deleted.\n• No-op actions: User actions fail silently or have no effect. For example, in invoiceninja/#11188,\nthe filter button in \"Customer > Documents\" does not sort or filter and always shows the full list.\n• Navigation failures: Pages fail to transition correctly. For example, in prestashop/#14796, a\nlogged-in user selecting any option in the back-office menu is redirected to the login page.",
          "bbox": [
            45.229000091552734,
            356.7392883300781,
            442.20611572265625,
            561.8283081054688
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 453
        }
      ],
      "identity_image": "assets/ref-02-identity.png"
    },
    {
      "number": 3,
      "key": "konstantinou2024actualexpected",
      "bib": {
        "author": "Konstantinou, Michael and Degiovanni, Renzo and Papadakis, Mike",
        "title": "Do LLMs Generate Test Oracles that Capture the Actual or the Expected Program Behaviour?",
        "year": "2024",
        "eprint": "2410.21136",
        "howpublished": "arXiv:2410.21136"
      },
      "url": "https://arxiv.org/abs/2410.21136",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/konstantinou-2024-actual-vs-expected-oracles/auto/konstantinou-2024-actual-vs-expected-oracles_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "f3140caad0d107492cfedc5b29e6c25957b7536889c8f78190224eae96b27a57",
      "pages": 12,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816120520",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文的受控研究確實發現，受測 LLM 較容易生成符合 actual behavior 的 oracle。本文 Introduction 的「tend to」合適；Related Work 的簡述應理解為該研究結果，並非所有模型與情境的定律。",
        "passages": [
          {
            "page": 2,
            "blocks": [
              2,
              4
            ],
            "label": "Introduction：actual 與 expected behavior 的研究發現",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2410.21136v1",
      "fulltext_url": "https://arxiv.org/pdf/2410.21136v1",
      "record_url": "https://arxiv.org/abs/2410.21136v1",
      "evidence": [
        {
          "id": "R03-E01",
          "number": 3,
          "page": 2,
          "label": "Introduction：actual 與 expected behavior 的研究發現",
          "image": "assets/ref-03-e01.png",
          "source_sha256": "f3140caad0d107492cfedc5b29e6c25957b7536889c8f78190224eae96b27a57",
          "image_sha256": "f74f01395ff390f0b966f27e9bca2b36ffd44692d748a2cb9f54c60460744021",
          "text": "Interestingly, our results show that LLMs are more likely to\ngenerate test oracles that capture the actual program behaviour\n(what is actually implemented) rather than the expected one,\ni.e., the intended behaviour. Additionally, we find that the\noverall performance of the LLMs is relatively low (less than\n50% accuracy) meaning that LLMs do not provide a strong\noracle correctness signal. Therefore, all LLMs suggestions will\nneed human inspection.\nTaken together, our results corroborate the conclusion that\nunless having meaningful test or variable names LLMs can\nmainly be used to capture the actual program behaviour (thus\nto be used for regression testing). Additionally, we find that\nLLMs could be a good addition to existing test generation\ntools, or to the test writing task, by using them to perform test\naugmentation. Overall, this work raises the awareness of the\npractical issues involved, advantages and disadvantages of the\nLLM-based test oracle generation abilities.",
          "bbox": [
            48.4640007019043,
            170.2575225830078,
            300.5218811035156,
            491.905029296875
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 709
        }
      ],
      "identity_image": "assets/ref-03-identity.png"
    },
    {
      "number": 4,
      "key": "bodicoat2025understanding",
      "bib": {
        "author": "Bodicoat, Adam and Jahangirova, Gunel and Terragni, Valerio",
        "title": "Understanding LLM-Driven Test Oracle Generation",
        "booktitle": "Proceedings of the 2nd IEEE/ACM International Conference on AI-Powered Software (AIware)",
        "year": "2025",
        "doi": "10.1109/AIWare69974.2025.00011",
        "pages": "29--39"
      },
      "url": "https://doi.org/10.1109/AIWare69974.2025.00011",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/bodicoat-2026-understanding-oracle-gen/auto/bodicoat-2026-understanding-oracle-gen_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "a5c315de17caf8c53860e31cbd5dd8baaf201d59d7ae1367c53c93ee499a192b",
      "pages": 11,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816113356",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要支持傳統回歸 oracle 依實作行為建立的背景；提示與 context 的效果由結果支持。對「actual rather than expected」的直接 LLM 實證，應連同 [3] 閱讀，不能把 [4] 的背景敘述誤當另一份相同實驗。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              5
            ],
            "label": "Abstract：依實作行為產生 regression oracle",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              17,
              18,
              19
            ],
            "label": "Introduction：context 與 prompt 的差異",
            "quote": null
          },
          {
            "page": 2,
            "blocks": [
              0,
              1
            ],
            "label": "主要發現：prompt 的影響大於模型選擇",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2601.05542v1",
      "fulltext_url": "https://arxiv.org/pdf/2601.05542v1",
      "record_url": "https://arxiv.org/abs/2601.05542v1",
      "evidence": [
        {
          "id": "R04-E01",
          "number": 4,
          "page": 1,
          "label": "Abstract：依實作行為產生 regression oracle",
          "image": "assets/ref-04-e01.png",
          "source_sha256": "a5c315de17caf8c53860e31cbd5dd8baaf201d59d7ae1367c53c93ee499a192b",
          "image_sha256": "b2abaf8c3ceca2cbae2797509581c7b5e2a66749080f120f1b1b4bcb44ceba36",
          "text": "Abstract—Automated unit test generation aims to improve\nsoftware quality while reducing the time and effort required for\ncreating tests manually. However, existing techniques primarily\ngenerate regression oracles that predicate on the implemented\nbehavior of the class under test. They do not address the oracle\nproblem: the challenge of distinguishing correct from incorrect\nprogram behavior.",
          "bbox": [
            48.4640007019043,
            182.9666290283203,
            300.833251953125,
            252.7775115966797
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 155
        },
        {
          "id": "R04-E02",
          "number": 4,
          "page": 1,
          "label": "Introduction：context 與 prompt 的差異",
          "image": "assets/ref-04-e02.png",
          "source_sha256": "a5c315de17caf8c53860e31cbd5dd8baaf201d59d7ae1367c53c93ee499a192b",
          "image_sha256": "714f64536260aa28a5529c7529ec3b2d80a205e226ad72e8f1102ecdfbaea752",
          "text": "①Oracles generated with more context compile and detect\nbugs more reliably. CUT-level context significantly outperforms\nother configurations, achieving 53.64% accuracy versus 40.74%\n(MUT) and 40.38% (test prefix only). This is an expected result.\n②Prompting style matters: zero-shot and few-shot prompts\nyield higher compilation rates (67.38% and 72.96%) and\naccuracy (54.56% and 51.30%) than CoT and ToT, which\nstruggle with low compilation (both below 50%).\n③Incorporating the CUT in the input prompt, along with\nzero-shot and few-shot prompting techniques, leads to the most\nconsistently accurate LLM-generated test oracles. However,\nour findings show there is potential for reasoning based prompt\ntechniques like CoT and ToT to be able to produce accurate\ntest oracles given their high accuracy when they do produce\ncompilable assertions.",
          "bbox": [
            311.14898681640625,
            492.49639892578125,
            565.2833251953125,
            680.070068359375
          ],
          "mode": "paragraph-crop",
          "width": 560,
          "height": 414
        },
        {
          "id": "R04-E03",
          "number": 4,
          "page": 2,
          "label": "主要發現：prompt 的影響大於模型選擇",
          "image": "assets/ref-04-e03.png",
          "source_sha256": "a5c315de17caf8c53860e31cbd5dd8baaf201d59d7ae1367c53c93ee499a192b",
          "image_sha256": "b1dea49fe92684672800132fcd30cee954aec11e399b1eec1dba0ece1bba0038",
          "text": "⑤Prompting strategy has a stronger impact on oracle\neffectiveness than LLM choice.\nThese findings suggest that prompt design and context\nplay a critical role in the effectiveness of LLM-based oracle\ngeneration.\nWhile reasoning-driven prompting (e.g., CoT,\nToT) shows potential when it compiles, zero-shot and few-shot\nprompting currently offer the best tradeoff between accuracy\nand robustness.\nOur study offers guidance for AI-assisted\ntesting tools usable by both testing and prompt experts in the\nFM era.",
          "bbox": [
            48.154998779296875,
            48.44737243652344,
            301.7625427246094,
            172.22506713867188
          ],
          "mode": "paragraph-crop",
          "width": 559,
          "height": 273
        }
      ],
      "identity_image": "assets/ref-04-identity.png"
    },
    {
      "number": 5,
      "key": "mughal2026sourceofauthority",
      "bib": {
        "author": "Mughal, Ali Hassaan and Bilal, Muhammad",
        "title": "LLM-Based Test Oracles: Source-of-Authority Taxonomy---A Systematic Literature Review",
        "year": "2026",
        "eprint": "2607.05031",
        "howpublished": "arXiv:2607.05031"
      },
      "url": "https://arxiv.org/abs/2607.05031",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/mughal-2026-llm-oracle-taxonomy-slr/auto/mughal-2026-llm-oracle-taxonomy-slr_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "794e0b25284208b5be6a7925b355f0a8bfd5565856ae9802a64bc617f50f878e",
      "pages": 21,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114609",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "本機來源為 arXiv v2，摘要明確記載 83 篇與 just over half。這是該回顧對納入研究的分類結果，不能推論成一半生成的 oracle 都錯誤。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              4
            ],
            "label": "Abstract：83 篇與逾半缺乏 specification",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2607.05031v2",
      "fulltext_url": "https://arxiv.org/pdf/2607.05031v2",
      "record_url": "https://arxiv.org/abs/2607.05031v2",
      "evidence": [
        {
          "id": "R05-E01",
          "number": 5,
          "page": 1,
          "label": "Abstract：83 篇與逾半缺乏 specification",
          "image": "assets/ref-05-e01.png",
          "source_sha256": "794e0b25284208b5be6a7925b355f0a8bfd5565856ae9802a64bc617f50f878e",
          "image_sha256": "d3aaf7d33a8c894b3a7ffe7a0424d29accf02f410aab873488c528a27d538989",
          "text": "Abstract—Large language models (LLMs) increasingly decide\nwhether software behaves correctly, either by writing a test oracle\nor by acting as one. Yet two oracles can look identical and rest on\ndifferent ground: one assertion encodes a written specification,\nanother only what the model learned in training. Prior secondary\nstudies sort oracles by form or by technique, rarely by the\nproperty that governs how far a verdict can be trusted: where its\nauthority comes from. This systematic literature review, reported\nunder the Preferred Reporting Items for Systematic Reviews\nand Meta-Analyses (PRISMA) 2020 guidelines, screens 2,436\nrecords to 54 included studies, extended by citation searching\n(snowballing) to 83 in total. We read the corpus along three\naxes: the source of an oracle’s authority, the form it takes, and\nthe mechanism that adjudicates it. Just over half of the corpus\nreaches a verdict with no specification at all. That is what lets\nthese oracles work on code with no specification to consult, and\nwhat leaves a challenged verdict with less to fall back on. Source\nand mechanism cross-cut rather than coincide, so a label such\nas LLM-as-a-judge names how a verdict is produced, not why\nit should be trusted. Oracle quality is most often judged by\nresemblance to a known oracle rather than by whether injected\nfaults are caught. The first question to ask of any LLM oracle is\ntherefore what one would point to in defending its verdict. The\nprotocol, search query, and per-study coding sheet are released.",
          "bbox": [
            48.46399688720703,
            176.6081085205078,
            300.5245056152344,
            415.71563720703125
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 527
        }
      ],
      "identity_image": "assets/ref-05-identity.png"
    },
    {
      "number": 6,
      "key": "gao2026guitester",
      "bib": {
        "author": "Gao, Yifei and Wu, Jiang and Chen, Xiaoyi and Yang, Yifan and Cui, Zhe and Ma, Tianyi and Zhang, Jiaming and Sang, Jitao",
        "title": "GUITester: Enabling GUI Agents for Exploratory Defect Discovery",
        "year": "2026",
        "eprint": "2601.04500",
        "howpublished": "arXiv:2601.04500"
      },
      "url": "https://arxiv.org/abs/2601.04500",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/guitester_2601.04500/auto/guitester_2601.04500_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "f19d1269435ad7b7b81781340567d9e02666ea6244ed0302d4b77d748b0c1829",
      "pages": 22,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816120819",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文直接命名 Goal-Oriented Masking 與 Execution-Bias Attribution，也提供含探索任務與缺陷類型的 GUITestBench。本文三次使用均能找到依據。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              11,
              12
            ],
            "label": "Introduction：兩種失敗模式的定義",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              5,
              6
            ],
            "label": "Abstract：143 個任務與 26 個缺陷",
            "quote": null
          },
          {
            "page": 4,
            "blocks": [
              3
            ],
            "label": "Benchmark：依觸發機制分類缺陷",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2601.04500v1",
      "fulltext_url": "https://arxiv.org/pdf/2601.04500v1",
      "record_url": "https://arxiv.org/abs/2601.04500v1",
      "evidence": [
        {
          "id": "R06-E01",
          "number": 6,
          "page": 1,
          "label": "Introduction：兩種失敗模式的定義",
          "image": "assets/ref-06-e01.png",
          "source_sha256": "f19d1269435ad7b7b81781340567d9e02666ea6244ed0302d4b77d748b0c1829",
          "image_sha256": "401124b7e2ebe53f71b097242837fd0d271e071c0f2e5f05ceaa34fb8ea68ba9",
          "text": "We identify two fundamental challenges that\nprevent existing GUI agents from effective ex-\nploratory testing: (i) Goal-Oriented Masking.\nMost GUI agents are optimized to maximize task\nsuccess rates, which inherently encourages robust-\nness against environmental obstacles. In a testing\ncontext, this goal-oriented nature leads the agent\nto perceive functional anomalies as traversable hur-\ndles rather than reportable defects. As shown in\nFigure 1(a), when encountering a non-responsive\nbutton, the agent’s policy autonomously seeks al-\nternative navigation paths to reach the goal. This\n“success-at-all-costs” behavior effectively masks\nthe defect, rendering it invisible to the quality as-\nsurance pipeline. (ii) Execution-Bias Attribution.\nExploratory testing lacks explicit oracles, requiring\nagents to distinguish between their own operational\nfailures (e.g., coordinate miscalculations) and gen-\nuine software defects. Due to the stochastic nature\nof MLLM interactions, current agents exhibit a sys-\ntematic bias toward self-attribution: erroneously\nassuming that any failure to trigger a state change\nstems from their own execution imprecision. As\nillustrated in Figure 1(b), GUI-Owl misinterprets a\nsystem-level rendering failure as a misaligned click,\ncausing the genuine defect to be misclassified as a\ntransient execution error in the logs.",
          "bbox": [
            304.1910095214844,
            303.3063049316406,
            526.8217163085938,
            667.5767822265625
          ],
          "mode": "paragraph-crop",
          "width": 491,
          "height": 802
        },
        {
          "id": "R06-E02",
          "number": 6,
          "page": 1,
          "label": "Abstract：143 個任務與 26 個缺陷",
          "image": "assets/ref-06-e02.png",
          "source_sha256": "f19d1269435ad7b7b81781340567d9e02666ea6244ed0302d4b77d748b0c1829",
          "image_sha256": "5904950545e4ea9b19e9fb3a2f66dc90738c58b671325bd16eb601a4bceac146",
          "text": "Exploratory GUI testing is essential for soft-\nware quality but suffers from high manual\ncosts.\nWhile Multi-modal Large Language\nModel (MLLM) agents excel in navigation,\nthey fail to autonomously discover defects due\nto two core challenges: Goal-Oriented Mask-\ning, where agents prioritize task completion\nover reporting anomalies, and Execution-Bias\nAttribution, where system defects are misiden-\ntified as agent errors. To address these, we first\nintroduce GUITestBench, the first interactive\nbenchmark for this task, featuring 143 tasks\nacross 26 defects. We then propose GUITester,\na multi-agent framework that decouples navi-\ngation from verification via two modules: (i) a\nPlanning-Execution Module (PEM) that proac-\ntively probes for defects via embedded testing\nintents, and (ii) a Hierarchical Reflection Mod-\nule (HRM) that resolves attribution ambiguity\nthrough interaction history analysis. GUITester\nachieves an F1-score of 48.90% (Pass@3) on\nGUITestBench, outperforming state-of-the-art\nbaselines (33.35%). Our work demonstrates\nthe feasibility of autonomous exploratory test-\ning and provides a robust foundation for future\nGUI quality assurance 1.",
          "bbox": [
            86.76599884033203,
            240.14633178710938,
            274.2816467285156,
            550.0640258789062
          ],
          "mode": "paragraph-crop",
          "width": 414,
          "height": 683
        },
        {
          "id": "R06-E03",
          "number": 6,
          "page": 4,
          "label": "Benchmark：依觸發機制分類缺陷",
          "image": "assets/ref-06-e03.png",
          "source_sha256": "f19d1269435ad7b7b81781340567d9e02666ea6244ed0302d4b77d748b0c1829",
          "image_sha256": "6f3c47ba4d8f53899c884633899590d945308626257f02136f18362ec2db27b9",
          "text": "5 diverse domains.\nUsing the exploratory task\nsynthesis strategies described above, we expand\nthese scenarios into 143 navigation tasks. The de-\ntailed distribution across defect types and appli-\ncation domains is shown in Figure 3. Based on\ndefect-triggering mechanisms, defects fall into two\ncategories: single-action defects, which are trig-\ngered by one action on a specific state (62.24%),\nand multi-action defects, which require a sequence\nof prerequisite actions (37.76%).",
          "bbox": [
            70.36599731445312,
            299.5903625488281,
            291.4414367675781,
            433.5247497558594
          ],
          "mode": "paragraph-crop",
          "width": 488,
          "height": 295
        }
      ],
      "identity_image": "assets/ref-06-identity.png"
    },
    {
      "number": 7,
      "key": "koong2026tuco",
      "bib": {
        "author": "Koong, Chorng-Shiuh and Lin, Yuan-Chun and Chen, Yao-Ting",
        "title": "Tuco: An LLM-driven prototype framework for web acceptance testing and E2E evidence generation",
        "booktitle": "Proceedings of the 22nd Taiwan Conference on Software Engineering (TCSE 2026)",
        "year": "2026",
        "address": "Taiwan"
      },
      "url": "",
      "source_path": "/home/ubuntu/tuco-lab/latex-writing/TCSE研討會投稿論文/Tuco An LLM-Driven Prototype Framework for Web Acceptance Testing and E2E Evidence Generation.pdf",
      "provenance": "author_manuscript",
      "download_url": "",
      "sha256": "9175c7711b805a02123098e61f4011f2da59ed66d87c8517cf877c1b37075346",
      "pages": 6,
      "pdf_metadata": {
        "format": "PDF 1.5",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "TeX",
        "producer": "MiKTeX pdfTeX-1.40.28",
        "creationDate": "D:20260614204624+08'00'",
        "modDate": "D:20260614204624+08'00'",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "本機 TCSE 作者原稿直接記載四階段產物、Gherkin Then assertion authority，以及 orchestrator–executor。這足以支持研究沿革與背景；不構成 JISE 本次實驗有執行四階段的證明。公開出版紀錄未在此項獨立核驗，來源明列為作者原稿。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              10
            ],
            "label": "Abstract：Tuco 原型與原先實驗",
            "quote": null
          },
          {
            "page": 2,
            "blocks": [
              8,
              9,
              10
            ],
            "label": "§III：四項產物與 Rule Extraction",
            "quote": null
          },
          {
            "page": 2,
            "blocks": [
              15,
              16,
              17,
              18
            ],
            "label": "§III.C：Gherkin 與 Then 斷言依據",
            "quote": null
          }
        ]
      },
      "source_version": "作者提供之 TCSE 原稿",
      "fulltext_url": "",
      "record_url": "",
      "evidence": [
        {
          "id": "R07-E01",
          "number": 7,
          "page": 1,
          "label": "Abstract：Tuco 原型與原先實驗",
          "image": "assets/ref-07-e01.png",
          "source_sha256": "9175c7711b805a02123098e61f4011f2da59ed66d87c8517cf877c1b37075346",
          "image_sha256": "c0a218f60c93ea2da675d9e7c714a79fd1cab1bdfdb4630e86478a1821935370",
          "text": "Abstract—This paper presents Tuco, a prototype framework\nthat uses large language models (LLMs) to support web accep-\ntance testing. In current software development, requirement-level\nacceptance still relies heavily on manual effort, while end-to-\nend (E2E) automated tests demand executable scripts whose\nassertions must reflect real requirements. Tuco constructs a\nstaged artifact pipeline that produces acceptance rules, web\noperation manuals, Gherkin scenarios, Playwright E2E tests,\nand operation recordings. Its browser execution stage adopts\nan orchestrator-executor design: the orchestrator dispatches\nscenarios and aggregates artifacts, while executor agents operate\nthe browser, evaluate observed behavior, and generate test\nscripts or candidate-defect evidence. A central design principle is\nGherkin-guided assertion strictness: generated Playwright asser-\ntions should verify the observable outcome stated in the Gherkin\nThen clause, and weakened assertions are reported as audit\nfindings rather than silently accepted as passing tests. In a post-fix\ncontrolled rerun over 63 scenarios across four modules, together\nwith an audit ablation over 12 pre-specified scenarios, Tuco\nsurfaced candidate defects related to data integrity, permission\nprotection, error handling, and specification mismatch. To verify\nwhether candidate items are real product defects and reduce\nmodel-verdict bias, three reviewers with about five years of\npractical experience retrospectively triaged the candidates under\na majority-vote rule; we report both Tuco-raw and reviewer-\nmajority confirmed results.",
          "bbox": [
            48.4640007019043,
            266.720947265625,
            300.5245666503906,
            525.7536010742188
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 571
        },
        {
          "id": "R07-E02",
          "number": 7,
          "page": 2,
          "label": "§III：四項產物與 Rule Extraction",
          "image": "assets/ref-07-e02.png",
          "source_sha256": "9175c7711b805a02123098e61f4011f2da59ed66d87c8517cf877c1b37075346",
          "image_sha256": "2550ff55d473a2f442aa0b5277e4acbd2367a48ef66f0469e43d3c1d12149924",
          "text": "Tuco organizes acceptance work around four artifacts: ex-\ntracted rules, a web operation manual, Gherkin scenarios,\nand browser-executed evidence. When the UI path, expected\nbehavior, or assertion target is underspecified, the browser\nexecution records the gap in an execution audit report that\nserves as input for orchestrator or human review.\nA. Rule Extraction\nRule extraction normalizes requirement documents into\nacceptance rules that can be cited during browser execu-\ntion. These documents may include feature descriptions, field\nconstraints, role permissions, error-handling rules, data-state\ntransitions, and exceptional cases. Tuco uses an LLM to\norganize these natural-language requirements while retaining\nthe source context for each rule.",
          "bbox": [
            48.463993072509766,
            539.410400390625,
            300.5218811035156,
            719.3519897460938
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 397
        },
        {
          "id": "R07-E03",
          "number": 7,
          "page": 2,
          "label": "§III.C：Gherkin 與 Then 斷言依據",
          "image": "assets/ref-07-e03.png",
          "source_sha256": "9175c7711b805a02123098e61f4011f2da59ed66d87c8517cf877c1b37075346",
          "image_sha256": "6e8d858c5b370ca598f631fcd1242f4fd4fd8f814a2275617933864fe2ce5102",
          "text": "C. Gherkin Acceptance Scenario Generation\nGherkin scenarios are drafted from the rules and operation\nmanual. The Given, When, and Then structure separates\nstate, action, and expected behavior.\nTuco uses Gherkin as the source for later Playwright as-\nsertions. Given describes the precondition, When describes\nthe user operation, and Then describes what the system must\nsatisfy. The Then clause becomes the basis for the generated\nPlaywright assertion.\nFor example, if a Then clause says that the system must\nreject a duplicate email address and display a clear error\nmessage, a later Playwright test cannot merely check that\nan API returned an error or that the page is still visible. It\nmust verify that the user-visible feedback appears and that the\nfeedback corresponds to the duplicate-email condition.",
          "bbox": [
            311.4779968261719,
            333.65411376953125,
            563.5357055664062,
            516.166015625
          ],
          "mode": "paragraph-crop",
          "width": 555,
          "height": 402
        }
      ],
      "identity_image": "assets/ref-07-identity.png"
    },
    {
      "number": 8,
      "key": "chevrot2025pinata",
      "bib": {
        "author": "Chevrot, Antoine and Vernotte, Alexandre and Falleri, Jean-R{\\'e}my and Blanc, Xavier and Legeard, Bruno and Cretin, Aymeric",
        "title": "Are Autonomous Web Agents Good Testers?",
        "journal": "Proceedings of the ACM on Software Engineering",
        "volume": "2",
        "number": "ISSTA",
        "year": "2025",
        "eprint": "2504.01495",
        "pages": "206--228",
        "doi": "10.1145/3728879"
      },
      "url": "https://doi.org/10.1145/3728879",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/chevrot-2025-autonomous-web-agents-testers/auto/chevrot-2025-autonomous-web-agents-testers_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "2e798f567557102966a76ca66fdb577bfe8b4f1abeef89c486b1ac53b53c8e82",
      "pages": 23,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816113219",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文有人工測例 benchmark，以及 AER／HER 的明確定義。使用的 arXiv v1 為五位作者，正式會議紀錄有六位作者（含 Aymeric Cretin）；這是版本差異，不能把第六位判為虛構。",
        "metadata_note": "截圖：arXiv:2504.01495v1（五位作者）；本文書目：正式 PACMSE／ISSTA 版本（六位作者）。官方會議作者名單已交叉確認。",
        "extra_links": [
          [
            "正式 ISSTA 2025 紀錄",
            "https://conf.researchr.org/details/issta-2025/issta-2025-papers/10/Are-Autonomous-Web-Agents-good-testers-"
          ]
        ],
        "passages": [
          {
            "page": 1,
            "blocks": [
              3
            ],
            "label": "Abstract：113 個人工測例",
            "quote": null
          },
          {
            "page": 13,
            "blocks": [
              1,
              2
            ],
            "label": "§4：Automation 與 Hallucination Error Rate",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2504.01495v1",
      "fulltext_url": "https://arxiv.org/pdf/2504.01495v1",
      "record_url": "https://arxiv.org/abs/2504.01495v1",
      "evidence": [
        {
          "id": "R08-E01",
          "number": 8,
          "page": 1,
          "label": "Abstract：113 個人工測例",
          "image": "assets/ref-08-e01.png",
          "source_sha256": "2e798f567557102966a76ca66fdb577bfe8b4f1abeef89c486b1ac53b53c8e82",
          "image_sha256": "5d7c1137a73bd04137acd93ebfddfb06cfbaae96790a5d2d96541e80ceeb7b5a",
          "text": "We contribute with (1) a benchmark of three offline web applications, and a suite of 113 manual test cases,\nsplit between passing and failing cases, to evaluate and compare ATAs performance, (2) SeeAct-ATA and\npinATA, two open-source ATA implementations capable of executing test steps, verifying assertions and\ngiving verdicts, and (3) comparative experiments using our benchmark that quantifies our ATAs effectiveness.\nFinally we also proceed to a qualitative evaluation to identify the limitations of PinATA, our best performing\nimplementation.",
          "bbox": [
            45.327999114990234,
            256.2188720703125,
            442.051513671875,
            320.98028564453125
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 144
        },
        {
          "id": "R08-E02",
          "number": 8,
          "page": 13,
          "label": "§4：Automation 與 Hallucination Error Rate",
          "image": "assets/ref-08-e02.png",
          "source_sha256": "2e798f567557102966a76ca66fdb577bfe8b4f1abeef89c486b1ac53b53c8e82",
          "image_sha256": "10c9dfc2b520c10a75bd49328ce597b28d82e48b36ff8b79538d9e3b2aa942bb",
          "text": "The AER rate describes the number of time the agent failed before the human without even\nhaving seen the failing step. These errors are likely due to the agent’s limitation in exploring the\nwebsite or validating an assertion.\nHallucinations are produced outputs that are coherent and grammatically correct but factually\nincorrect or nonsensical. In that case, hallucinations would be unlawfully validated step assertions\nand the HER quantify them.",
          "bbox": [
            44.959999084472656,
            86.17494201660156,
            440.93206787109375,
            156.91354370117188
          ],
          "mode": "paragraph-crop",
          "width": 873,
          "height": 157
        }
      ],
      "identity_image": "assets/ref-08-identity.png"
    },
    {
      "number": 9,
      "key": "shahbandeh2024naviqate",
      "bib": {
        "author": "Shahbandeh, Mobina and Alian, Parsa and Nashid, Noor and Mesbah, Ali",
        "title": "NaviQAte: Functionality-Guided Web Application Navigation",
        "year": "2024",
        "eprint": "2409.10741",
        "howpublished": "arXiv:2409.10741"
      },
      "url": "https://arxiv.org/abs/2409.10741",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/shahbandeh-2024-naviqate/auto/shahbandeh-2024-naviqate_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "3e98086d14e21fc896097a93ddc4888bafbff3c8c2920b6751be115d5892869d",
      "pages": 21,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114335",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要直接將 web exploration 描述為 question-and-answer task，與本文句子一致。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              2
            ],
            "label": "Abstract：question-and-answer exploration",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2409.10741v1",
      "fulltext_url": "https://arxiv.org/pdf/2409.10741v1",
      "record_url": "https://arxiv.org/abs/2409.10741v1",
      "evidence": [
        {
          "id": "R09-E01",
          "number": 9,
          "page": 1,
          "label": "Abstract：question-and-answer exploration",
          "image": "assets/ref-09-e01.png",
          "source_sha256": "3e98086d14e21fc896097a93ddc4888bafbff3c8c2920b6751be115d5892869d",
          "image_sha256": "93b7c56ca1c69bfeb5d14b1e2d52f6e7add8cf00ce2742c4a52d6a032ec2b8b2",
          "text": "End-to-end web testing is challenging due to the need to explore diverse web application functionalities.\nCurrent state-of-the-art methods, such as WebCanvas, are not designed for broad functionality exploration;\nthey rely on specific, detailed task descriptions, limiting their adaptability in dynamic web environments. We\nintroduce NaviQAte, which frames web application exploration as a question-and-answer task, generating\naction sequences for functionalities without requiring detailed parameters. Our three-phase approach utilizes\nadvanced large language models like GPT-4o for complex decision-making and cost-effective models, such\nas GPT-4o mini, for simpler tasks. NaviQAte focuses on functionality-guided web application navigation,\nintegrating multi-modal inputs such as text and images to enhance contextual understanding. Evaluations on\nthe Mind2Web-Live and Mind2Web-Live-Abstracted datasets show that NaviQAte achieves a 44.23% success\nrate in user task navigation and a 38.46% success rate in functionality navigation, representing a 15% and\n33% improvement over WebCanvas. These results underscore the effectiveness of our approach in advancing\nautomated web application testing.",
          "bbox": [
            45.327999114990234,
            166.8824462890625,
            442.05426025390625,
            297.4637756347656
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 288
        }
      ],
      "identity_image": "assets/ref-09-identity.png"
    },
    {
      "number": 10,
      "key": "mesbah2012crawljax",
      "bib": {
        "author": "Mesbah, Ali and van Deursen, Arie and Lenselink, Stefan",
        "title": "Crawling Ajax-based web applications through dynamic analysis of user interface state changes",
        "journal": "ACM Transactions on the Web",
        "volume": "6",
        "number": "1",
        "pages": "3",
        "year": "2012",
        "doi": "10.1145/2109205.2109208"
      },
      "url": "https://doi.org/10.1145/2109205.2109208",
      "source_path": "/home/ubuntu/tuco-lab/tmp/citation-audit/downloads/mesbah2012crawljax.pdf",
      "provenance": "downloaded_primary",
      "download_url": "https://people.ece.ubc.ca/amesbah/resources/papers/tweb-final.pdf",
      "sha256": "ce355ef34f933f170cc7391086ce87f5c972fd8a38f4ba7e91b3a78ce6e928d9",
      "pages": 30,
      "pdf_metadata": {
        "format": "PDF 1.5",
        "title": "Crawling Ajax-Based Web Applications through Dynamic Analysis of User Interface State Changes",
        "author": "Ali Mesbah, Arie van Deursen, Stefan Lenselink",
        "subject": "",
        "keywords": "Ajax, Crawling, DOM crawling, Web 2.0, dynamic analysis, hidden web",
        "creator": "dvips(k) 5.98 Copyright 2009 Radical Eye Software",
        "producer": "Acrobat Distiller 8.1.0 (Windows)",
        "creationDate": "D:20120229223905+08'00'",
        "modDate": "D:20120229223935+08'",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "Crawljax 原文支持它是以 GUI 狀態變化驅動的動態爬蟲。Temac 超越該基線的比較證據來自 [11]，不是這篇 2012 年論文。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              3
            ],
            "label": "Abstract：動態分析 GUI 狀態變化",
            "quote": "This article describes a novel technique for crawling AJAX-based applications through automatic dynamic analysis of user-interface-state changes in Web browsers."
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://people.ece.ubc.ca/amesbah/resources/papers/tweb-final.pdf",
      "record_url": "https://doi.org/10.1145/2109205.2109208",
      "evidence": [
        {
          "id": "R10-E01",
          "number": 10,
          "page": 1,
          "label": "Abstract：動態分析 GUI 狀態變化",
          "image": "assets/ref-10-e01.png",
          "source_sha256": "ce355ef34f933f170cc7391086ce87f5c972fd8a38f4ba7e91b3a78ce6e928d9",
          "image_sha256": "449537fc938fa62e4b03090179f3e26dc0fe188a846f2dfecc39e8465265bf0c",
          "text": "This article describes a novel technique for crawling AJAX-based applications through automatic dynamic analysis of user-interface-state changes in Web browsers.",
          "bbox": [
            106.83999633789062,
            201.41946411132812,
            505.286865234375,
            233.30966186523438
          ],
          "rects": [
            [
              313.5426940917969,
              203.41946411132812,
              503.1838073730469,
              211.38955688476562
            ],
            [
              108.84024047851562,
              213.37936401367188,
              503.286865234375,
              221.34976196289062
            ],
            [
              108.83999633789062,
              223.33956909179688,
              143.81918334960938,
              231.30966186523438
            ]
          ],
          "mode": "short-quote-lines",
          "width": 997,
          "height": 81
        }
      ]
    },
    {
      "number": 11,
      "key": "liu2025temac",
      "bib": {
        "author": "Liu, Chenxu and Gu, Zhiyu and Wu, Guoquan and Zhang, Ying and Wei, Jun and Xie, Tao",
        "title": "Temac: Multi-Agent Collaboration for Automated Web GUI Testing",
        "year": "2025",
        "eprint": "2506.00520",
        "howpublished": "arXiv:2506.00520"
      },
      "url": "https://arxiv.org/abs/2506.00520",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/liu-2025-temac/auto/liu-2025-temac_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "91392fb6446e63cb90e62722fda2a60a4a6a283c0dcb0b3d913986ce660e000a",
      "pages": 12,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114017",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要說明先以既有 AWGT 探索、再由 LLM agents 執行未覆蓋功能；§4 的 baseline 列表含 Crawljax。因此可支持這個比較方向，不能擴大成 Temac 在每個功能上都勝出。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              6,
              7
            ],
            "label": "Abstract：執行未覆蓋功能及評估結果",
            "quote": null
          },
          {
            "page": 7,
            "blocks": [
              11,
              12
            ],
            "label": "評估基線：Crawljax 與 FragGen",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2506.00520v1",
      "fulltext_url": "https://arxiv.org/pdf/2506.00520v1",
      "record_url": "https://arxiv.org/abs/2506.00520v1",
      "evidence": [
        {
          "id": "R11-E01",
          "number": 11,
          "page": 1,
          "label": "Abstract：執行未覆蓋功能及評估結果",
          "image": "assets/ref-11-e01.png",
          "source_sha256": "91392fb6446e63cb90e62722fda2a60a4a6a283c0dcb0b3d913986ce660e000a",
          "image_sha256": "7fadbcd908f046e253d7a9ad712075ae932b225f3bf0100b3dc0be45e628a20b",
          "text": "To address these challenges, in this paper, we propose Temac,\nan approach that enhances automated web GUI testing using\nLLM-based multi-agent collaboration, aiming to maintain both\nexploration breadth and depth to increase code coverage. Temac\nis motivated by our insight that LLMs can enhance automated\nweb GUI testing in executing complex functionalities, while the\ninformation discovered during automated web GUI testing can, in\nturn, be provided as the domain knowledge to improve the success\nrate of LLM-based task planning and execution. Specifically,\ngiven a web application, Temac initially runs an existing approach\nof automated web GUI testing to broadly explore application\nstates. When the testing coverage stagnates, Temac then employs\nLLM-based agents to summarize the collected multi-modal in-\nformation into a structured and concise knowledge base and to\ninfer not-covered functionalities. Guided by this knowledge base,\nTemac finally uses specialized LLM-based agents to target and\nexecute these not-covered functionalities, reaching deeper states\nbeyond those explored by a testing approach without using LLMs.\nOur evaluation results show that Temac improves state-of-the-\nart approaches of automated web GUI testing from 12.5% to\n60.3% on average code coverage on six complex open-source\nweb applications, while revealing 445 unique faults in the top 20\nreal-world web applications. These results strongly demonstrate\nthe effectiveness and the general applicability of Temac.",
          "bbox": [
            48.463958740234375,
            356.4712219238281,
            300.5245056152344,
            595.6648559570312
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 527
        },
        {
          "id": "R11-E02",
          "number": 11,
          "page": 7,
          "label": "評估基線：Crawljax 與 FragGen",
          "image": "assets/ref-11-e02.png",
          "source_sha256": "91392fb6446e63cb90e62722fda2a60a4a6a283c0dcb0b3d913986ce660e000a",
          "image_sha256": "732695550b6137348a7ab2d6f0627d08fcf50fa938cadae84f27f420c7dde47d",
          "text": "• Crawljax [3]. A model-based AWGT approach that is\nwidely used by existing work [5]–[7], [45] as a baseline.\n• FragGen [6]. A model-based AWGT approach that is\ndeveloped based on Crawljax, equipped with an effective\nstate abstraction approach using screenshot matching.",
          "bbox": [
            58.42702102661133,
            419.76751708984375,
            300.5187072753906,
            478.55108642578125
          ],
          "mode": "paragraph-crop",
          "width": 534,
          "height": 130
        }
      ],
      "identity_image": "assets/ref-11-identity.png"
    },
    {
      "number": 12,
      "key": "hu2024auitestagent",
      "bib": {
        "author": "Hu, Yongxiang and Wang, Xuan and Wang, Yingchuan and Zhang, Yu and Guo, Shiyu and Chen, Chaoyi and Wang, Xin and Zhou, Yangfan",
        "title": "AUITestAgent: Automatic Requirements Oriented GUI Function Testing",
        "year": "2024",
        "eprint": "2407.09018",
        "howpublished": "arXiv:2407.09018"
      },
      "url": "https://arxiv.org/abs/2407.09018",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/hu-2024-auitestagent/auto/hu-2024-auitestagent_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "0ad43a7195b07faa5dfdc922407cd69fcd27be638f84510535c9c9f5305249a8",
      "pages": 10,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114205",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要直接說明對 mobile apps 依自然語言 requirements 執行 GUI interaction 與 function verification。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              27
            ],
            "label": "Abstract：行動應用的需求導向 GUI 測試",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2407.09018v1",
      "fulltext_url": "https://arxiv.org/pdf/2407.09018v1",
      "record_url": "https://arxiv.org/abs/2407.09018v1",
      "evidence": [
        {
          "id": "R12-E01",
          "number": 12,
          "page": 1,
          "label": "Abstract：行動應用的需求導向 GUI 測試",
          "image": "assets/ref-12-e01.png",
          "source_sha256": "0ad43a7195b07faa5dfdc922407cd69fcd27be638f84510535c9c9f5305249a8",
          "image_sha256": "f1cccb830c19a9078027863e3e2ab407bf0f287dfbf7cc85edbcdcd62c8b8ab8",
          "text": "The Graphical User Interface (GUI) is how users interact with mo-\nbile apps. To ensure it functions properly, testing engineers have to\nmake sure it functions as intended, based on test requirements that\nare typically written in natural language. While widely adopted\nmanual testing and script-based methods are effective, they demand\nsubstantial effort due to the vast number of GUI pages and rapid\niterations in modern mobile apps. This paper introduces AUITestA-\ngent, the first automatic, natural language-driven GUI testing tool\nfor mobile apps, capable of fully automating the entire process of\nGUI interaction and function verification. Since test requirements\ntypically contain interaction commands and verification oracles.\nAUITestAgent can extract GUI interactions from test requirements\nvia dynamically organized agents. Then, AUITestAgent employs a\nmulti-dimensional data extraction strategy to retrieve data relevant\nto the test requirements from the interaction trace and perform\nverification. Experiments on customized benchmarks 1 demonstrate\nthat AUITestAgent outperforms existing tools in the quality of gen-\nerated GUI interactions and achieved the accuracy of verifications of\n94%. Moreover, field deployment in Meituan has shown AUITestA-\ngent’s practical usability, with it detecting 4 new functional bugs\nduring 10 regression tests in two months.",
          "bbox": [
            53.02000045776367,
            330.415771484375,
            296.06072998046875,
            559.5568237304688
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 506
        }
      ],
      "identity_image": "assets/ref-12-identity.png"
    },
    {
      "number": 13,
      "key": "kolthoff2025guispector",
      "bib": {
        "author": "Kolthoff, Kristian and Kretzer, Felix and Ponzetto, Simone Paolo and Maedche, Alexander and Bartelt, Christian",
        "title": "GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes",
        "year": "2025",
        "eprint": "2510.04791",
        "howpublished": "arXiv:2510.04791"
      },
      "url": "https://arxiv.org/abs/2510.04791",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/kolthoff-2025-guispector/auto/kolthoff-2025-guispector_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "80157cb55cd626d9f048675691b8a6b4797844038e368f8731053ef6707c03fb",
      "pages": 4,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114850",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要直接說明在 GUI prototypes 驗證 natural-language requirements，與本文引用用途一致。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              17
            ],
            "label": "Abstract：GUI prototype 的需求驗證",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2510.04791v1",
      "fulltext_url": "https://arxiv.org/pdf/2510.04791v1",
      "record_url": "https://arxiv.org/abs/2510.04791v1",
      "evidence": [
        {
          "id": "R13-E01",
          "number": 13,
          "page": 1,
          "label": "Abstract：GUI prototype 的需求驗證",
          "image": "assets/ref-13-e01.png",
          "source_sha256": "80157cb55cd626d9f048675691b8a6b4797844038e368f8731053ef6707c03fb",
          "image_sha256": "d79c6eaa05c137a8aef85d2bc162b17739d89853031c51c01d38e58ef01fc0e2",
          "text": "Graphical user interfaces (GUIs) are foundational to interactive\nsystems and play a pivotal role in early requirements elicitation\nthrough prototyping. Ensuring that GUI implementations fulfill\nnatural language (NL) requirements is essential for robust software\nengineering, especially as LLM-driven programming agents become\nincreasingly integrated into development workflows. Existing GUI\ntesting approaches, whether traditional or LLM-driven, often fall\nshort in handling the complexity of modern interfaces, and typically\nlack actionable feedback and effective integration with automated\ndevelopment agents. In this paper, we introduce GUISpector, a novel\nframework that leverages a multi-modal (M)LLM-based agent for\nthe automated verification of NL requirements in GUI prototypes.\nFirst, GUISpector adapts a MLLM agent to interpret and opera-\ntionalize NL requirements, enabling to autonomously plan and\nexecute verification trajectories across GUI applications. Second,\nGUISpector systematically extracts detailed NL feedback from the\nagent’s verification process, providing developers with actionable\ninsights that can be used to iteratively refine the GUI artifact or\ndirectly inform LLM-based code generation in a closed feedback\nloop. Third, we present an integrated tool that unifies these capabil-\nities, offering practitioners an accessible interface for supervising\nverification runs, inspecting agent rationales and managing the\nend-to-end requirements verification process. We evaluated GUIS-\npector on a comprehensive set of 150 requirements based on 900\nacceptance criteria annotations across diverse GUI applications,\ndemonstrating effective detection of requirement satisfaction and\nviolations and highlighting its potential for seamless integration\nof actionable feedback into automated LLM-driven development\nworkflows. The video presentation of GUISpector is available at:\nhttps://youtu.be/JByYF6BNQeE, showcasing its main capabilities.",
          "bbox": [
            52.96699905395508,
            283.1184387207031,
            296.05670166015625,
            610.9575805664062
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 723
        }
      ],
      "identity_image": "assets/ref-13-identity.png"
    },
    {
      "number": 14,
      "key": "kong2026webtestbench",
      "bib": {
        "author": "Kong, Fanheng and Zhang, Jingyuan and Yue, Yang and Sun, Chenxi and Tian, Yang and Feng, Shi and Yang, Xiaocui and Wang, Daling and Tian, Yu and Du, Jun and Zeng, Wenchong and Li, Han and Gai, Kun",
        "title": "WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing",
        "year": "2026",
        "eprint": "2603.25226",
        "howpublished": "arXiv:2603.25226"
      },
      "url": "https://arxiv.org/abs/2603.25226",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/webtestbench_2603.25226/auto/webtestbench_2603.25226_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
      "pages": 24,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816112726",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "Table 2 支持所有受測模型 end-to-end F1 < 30%，§5 說明 precision／recall 的取捨；§3.2 描述用 Lovable 合成並迭代應用，附錄表格列 1,750 測項與 448 個 Fail。448 的單位是失敗測項，不是已去重的獨立 bug。",
        "passages": [
          {
            "page": 6,
            "blocks": [
              0,
              1,
              2,
              3
            ],
            "label": "Table 2：完整模型結果",
            "quote": null
          },
          {
            "page": 7,
            "blocks": [
              4
            ],
            "label": "Overall Performance：precision／recall 取捨",
            "quote": null
          },
          {
            "page": 4,
            "blocks": [
              15
            ],
            "label": "§3.2：AI 生成應用",
            "quote": null
          },
          {
            "page": 4,
            "blocks": [
              18
            ],
            "label": "§3.2：測項標註與迭代生成",
            "quote": null
          },
          {
            "page": 12,
            "blocks": [
              9
            ],
            "label": "附錄統計：1,750 total items、1,302／448 pass／fail",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2603.25226v1",
      "fulltext_url": "https://arxiv.org/pdf/2603.25226v1",
      "record_url": "https://arxiv.org/abs/2603.25226v1",
      "evidence": [
        {
          "id": "R14-E01",
          "number": 14,
          "page": 6,
          "label": "Table 2：完整模型結果",
          "image": "assets/ref-14-e01.png",
          "source_sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
          "image_sha256": "6171ebdb63baaf28f3442dc51e8c27f7d09d5f2e8dcc8d8a98ba33e658dfad63",
          "text": "Model\n#Turns\n#Tokens\nFunctionality\nConstraint\nInteraction\nContent\nOverall\nCov.\nF1\nCov.\nF1\nCov.\nF1\nCov.\nF1\nCov.\nP\nR\nF1\nOpen-Source LLMs\nMinimax-M2.1\n41.7\n3.58M\n77.9\n12.3\n40.4\n15.8\n42.2\n19.9\n47.1\n7.7\n60.1\n22.3\n14.6\n15.2\nQwen3-Coder-Next\n63.4\n6.24M\n77.6\n14.1\n48.3\n23.8\n42.7\n11.4\n35.9\n4.3\n60.4\n27.8\n15.8\n17.3\nGLM-4.7\n41.6\n3.47M\n79.9\n16.5\n47.3\n20.5\n36.0\n17.2\n48.9\n4.3\n61.1\n26.7\n16.6\n18.1\nGLM-5\n41.3\n3.71M\n79.7\n11.9\n50.1\n26.9\n41.4\n20.9\n50.6\n3.4\n63.1\n30.4\n15.6\n19.0\nStep-3.5-Flash\n57.0\n3.37M\n79.8\n20.1\n53.6\n27.9\n48.5\n21.2\n60.6\n2.6\n66.0\n34.6\n20.8\n23.4\nMiMo-V2-Flash\n59.8\n7.26M\n80.0\n21.9\n48.7\n29.2\n48.0\n20.3\n50.3\n7.3\n63.5\n34.8\n24.6\n25.1\nClosed-Source LLMs\nClaude Opus 4.5\n42.9\n2.60M\n83.3\n18.8\n50.3\n21.2\n40.8\n14.7\n42.1\n6.8\n63.2\n33.0\n16.5\n20.2\nClaude Sonnet 4.5\n37.6\n1.90M\n81.0\n22.2\n47.7\n22.5\n46.7\n19.9\n51.6\n1.7\n63.7\n32.1\n19.7\n21.9\nGPT-5.2\n69.5\n7.43M\n76.9\n25.3\n51.9\n21.5\n43.0\n23.2\n46.1\n6.2\n61.0\n24.7\n25.2\n22.9\nGPT-5.1\n30.3\n0.87M\n76.4\n30.9\n51.2\n26.9\n49.7\n22.0\n57.5\n15.3\n63.1\n25.8\n33.3\n26.4\nTable 2: Web testing performance of representative LLM on WebTestBench under the WebTester framework. We\nreport the Coverage metric (Cov.) for checklist generation, and the average number of iteration turns (#Turns),\naverage context tokens (#Tokens) per instance, and Precision/Recall/F1 metrics for defect detection. Results in bold\nand underline denote the best and second-best performances.",
          "bbox": [
            70.05699920654297,
            74.50076293945312,
            526.149658203125,
            290.8200988769531
          ],
          "mode": "paragraph-crop",
          "width": 1004,
          "height": 477
        },
        {
          "id": "R14-E02",
          "number": 14,
          "page": 7,
          "label": "Overall Performance：precision／recall 取捨",
          "image": "assets/ref-14-e02.png",
          "source_sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
          "image_sha256": "f3d42440e527304c54cedc0df5f0bc57eacf5c9e4cf48eaa42d055abf6de3a0a",
          "text": "updates are easily mistaken as functional failures,\nreflecting insufficient model understanding of dy-\nnamic web behavior. On the other hand, the low\nrecall poses a greater risk. Most models fail to ex-\nceed a 25% recall, meaning the major real defects\nremain undetected. Beyond limited defect cogni-\ntion, these false-negatives are partly from a default-\ncorrectness bias, where models default to a pass\njudgment when no explicit evidence is observed.\nAdditionally, we observe a strategic divergence in\nCUA behavior: they either employ an aggressive\ndetection strategy that favors recall at the cost of\nprecision (e.g., GPT-5.1 achieves 33.3% recall but\nonly 25.8% precision) or adopt a conservative one\nthat prioritizes precision while overlooking numer-\nous real defects (e.g., MiMo-V2-Flash achieves\n34.8% precision but only 24.6% recall).\nLong-horizon Interaction Unreliability. Com-\npleting a comprehensive web defect detection pro-\ncess typically requires dozens of interaction turns\nand millions of tokens. For example, Step-3.5-\nFlash requires an average of 57.0 turns and 3.37M\ntokens per sample. Such long-horizon tasks de-\nmand rigorous long-context memory and planning\nstability. Specifically, as the interaction history\naccumulates, models become increasingly suscepti-\nble to tracking failures, resulting in the loss of prior\nstates or the execution of redundant operations.",
          "bbox": [
            69.9729995727539,
            249.47177124023438,
            291.5415344238281,
            627.6577758789062
          ],
          "mode": "paragraph-crop",
          "width": 489,
          "height": 833
        },
        {
          "id": "R14-E03",
          "number": 14,
          "page": 4,
          "label": "§3.2：AI 生成應用",
          "image": "assets/ref-14-e03.png",
          "source_sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
          "image_sha256": "87f35ec4eaff296df5eaf87f1dbf7fb4628f143eee8f2c6f02137d684517525c",
          "text": "ing agents requires environments that are ecolog-\nically valid and contain diverse and non-trivial\ndefects.\nHowever, standard web resources of-\nten present limitations. Commercial websites are\ntypically well tested and continuously updated,\nwhich makes them unsuitable as benchmark sam-\nples. Open-source projects often feature simple de-\nsigns with shallow structures or limited interactive\nfunctionality, and therefore fail to reflect realistic\nuser interactions. To bridge this gap, we utilize\nLovable.dev1, an AI-powered web development\nplatform that generates complete websites from\nuser instructions, to synthesize web application\nprojects. Through this process, we obtain an initial\nweb application for each instruction, providing a\nrealistic webpages for defect detection.\nGold Checklist and Result Annotation. Given an\ndevelopment instruction and its application, human\nannotators construct a testable checklist. Inspired\nby software quality and evaluation standards (e.g.,\nISO/IEC 25010 (ISO/IEC, 2023)), and adapting",
          "bbox": [
            69.9729995727539,
            472.7473449707031,
            291.4489440917969,
            755.9600219726562
          ],
          "mode": "paragraph-crop",
          "width": 489,
          "height": 624
        },
        {
          "id": "R14-E04",
          "number": 14,
          "page": 4,
          "label": "§3.2：測項標註與迭代生成",
          "image": "assets/ref-14-e04.png",
          "source_sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
          "image_sha256": "ac0f4c607173377f0a1b16545df771bad25ee06be8a0046ae48e044114903330",
          "text": "Annotators derive a checklist from the instruc-\ntion, and then interact with the actual application to\nalign test cases with the implemented components\nand interaction flows. Finally, annotators execute\nthe checklist by interacting with the website and\ndocument the Pass/Fail status of each item and pro-\nvide concise bug reports for failures.\nIterative Refinement. Initial synthesis often pro-\nduces applications with few defects, which limits\nthe effectiveness in discriminatively evaluating web\ntesting capabilities. To address this, annotators per-\nform a iterative refinement. This involves revising\nthe instruction for re-generate app or continuing\nthe conversation with lovable.dev to add new fea-\ntures. Throughout this iteration, the checklist and\nresults are updated synchronously until the samples\ncontain sufficient defects for evaluation.\nQuality Control. To ensure the quality of the\ndataset, all annotators undergo related training and\nconduct cross-validation during the annotation pro-\ncess.\nFinally, a senior annotation leader (non-\nauthors) performs a final scan of the entire dataset,\nproviding feedback and guiding annotators to op-",
          "bbox": [
            305.3689880371094,
            461.2893371582031,
            526.72509765625,
            775.9959716796875
          ],
          "mode": "paragraph-crop",
          "width": 488,
          "height": 694
        },
        {
          "id": "R14-E05",
          "number": 14,
          "page": 12,
          "label": "附錄統計：1,750 total items、1,302／448 pass／fail",
          "image": "assets/ref-14-e05.png",
          "source_sha256": "dbe4d01bd4d65580f620f10015fd9cd98b1c91984f5c2e7931ae00dfca1631eb",
          "image_sha256": "bd768c2d4c577a2f5936457c568c02747819fbd3fbf49d924ced6e600c28ebab",
          "text": "Total items\n1750\nFunctionality\n854\nConstraint\n398\nInteraction\n247\nContent\n251\nTotal Pass / Fail items\n1302/448\nFunctionality (Pass / Fail)\n653/201\nConstraint (Pass / Fail)\n270/128\nInteraction (Pass / Fail)\n176/71\nContent (Pass / Fail)\n203/48",
          "bbox": [
            340.00177001953125,
            271.054931640625,
            499.41571044921875,
            379.9933166503906
          ],
          "mode": "paragraph-crop",
          "width": 351,
          "height": 240
        }
      ],
      "identity_image": "assets/ref-14-identity.png"
    },
    {
      "number": 15,
      "key": "barr2015oracle",
      "bib": {
        "author": "Barr, Earl T. and Harman, Mark and McMinn, Phil and Shahbaz, Muzammil and Yoo, Shin",
        "title": "The oracle problem in software testing: A survey",
        "journal": "IEEE Transactions on Software Engineering",
        "volume": "41",
        "number": "5",
        "pages": "507--525",
        "year": "2015",
        "doi": "10.1109/TSE.2014.2372785"
      },
      "url": "https://doi.org/10.1109/TSE.2014.2372785",
      "source_path": "/home/ubuntu/tuco-lab/tmp/citation-audit/downloads/barr2015oracle.pdf",
      "provenance": "downloaded_primary",
      "download_url": "https://philmcminn.com/publications/barr2015.pdf",
      "sha256": "6c5689497594a75f08c1b37b3dd361e11c54946bbbf79f670d660f5d298d3d51",
      "pages": 31,
      "pdf_metadata": {
        "format": "PDF 1.5",
        "title": "The Oracle Problem in Software Testing: A Survey",
        "author": "Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, Shin Yoo",
        "subject": "",
        "keywords": "",
        "creator": "LaTeX with hyperref package",
        "producer": "pdfTeX-1.40.13",
        "creationDate": "D:20141117090220Z",
        "modDate": "D:20141117090220Z",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "這篇 survey 將 test oracle 一詞追溯至 1978 年，足以支持 oracle problem 早於當代 LLM 的背景句。",
        "passages": [
          {
            "page": 5,
            "blocks": [
              23
            ],
            "label": "§3：test oracle 的歷史",
            "quote": "The term “test oracle” first appeared in William Howden’s seminal work in 1978 [99]."
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://philmcminn.com/publications/barr2015.pdf",
      "record_url": "https://doi.org/10.1109/TSE.2014.2372785",
      "evidence": [
        {
          "id": "R15-E01",
          "number": 15,
          "page": 5,
          "label": "§3：test oracle 的歷史",
          "image": "assets/ref-15-e01.png",
          "source_sha256": "6c5689497594a75f08c1b37b3dd361e11c54946bbbf79f670d660f5d298d3d51",
          "image_sha256": "f6677c77f108498fc5f644a1b36da73c528841b18cd589ae1842e7a7b6f425cc",
          "text": "The term “test oracle” first appeared in William Howden’s seminal work in 1978 [99].",
          "bbox": [
            309.97796630859375,
            544.6494140625,
            522.1964111328125,
            570.5670166015625
          ],
          "rects": [
            [
              311.97796630859375,
              546.6494140625,
              520.1964111328125,
              556.6119995117188
            ],
            [
              311.97796630859375,
              558.6044311523438,
              485.755615234375,
              568.5670166015625
            ]
          ],
          "mode": "short-quote-lines",
          "width": 532,
          "height": 66
        }
      ]
    },
    {
      "number": 16,
      "key": "hossain2025togll",
      "bib": {
        "author": "Hossain, Soneya Binta and Dwyer, Matthew B.",
        "title": "TOGLL: Correct and Strong Test Oracle Generation with LLMs",
        "booktitle": "Proceedings of the 47th IEEE/ACM International Conference on Software Engineering (ICSE)",
        "year": "2025",
        "doi": "10.1109/ICSE55347.2025.00098",
        "eprint": "2405.03786",
        "pages": "1475--1487"
      },
      "url": "https://doi.org/10.1109/ICSE55347.2025.00098",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/hossain-2024-togll/auto/hossain-2024-togll_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "25c607fa39bb553e4e6a3387a49fa6a478a3087eea202534f54858d9da5b169f",
      "pages": 13,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816115319",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "Abstract 與 RQ2 都直接報告 3.8 倍 assertion oracles、4.9 倍 exception oracles。本文正確取用前者，指的是正確 oracle 數量，並非準確率增加 3.8 倍。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              8
            ],
            "label": "Abstract：與 TOGA 的比較",
            "quote": null
          },
          {
            "page": 7,
            "blocks": [
              41
            ],
            "label": "RQ2 Finding：3.8 倍與 4.9 倍",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2405.03786v2",
      "fulltext_url": "https://arxiv.org/pdf/2405.03786v2",
      "record_url": "https://arxiv.org/abs/2405.03786v2",
      "evidence": [
        {
          "id": "R16-E01",
          "number": 16,
          "page": 1,
          "label": "Abstract：與 TOGA 的比較",
          "image": "assets/ref-16-e01.png",
          "source_sha256": "25c607fa39bb553e4e6a3387a49fa6a478a3087eea202534f54858d9da5b169f",
          "image_sha256": "3b12c050c1f3d6b5767f77eb44fad79f2abc013ce49009f591248edad62851f4",
          "text": "In this research, we present the first comprehensive study\nto investigate the capabilities of LLMs in generating correct,\ndiverse, and strong test oracles capable of effectively identifying\na large number of unique bugs. To this end, we fine-tuned\nseven code LLMs using six distinct prompts on a large dataset\nconsisting of 110 Java projects. Utilizing the most effective fine-\ntuned LLM and prompt pair, we introduce TOGLL, a novel\nLLM-based method for test oracle generation. To investigate the\ngeneralizability of TOGLL, we conduct studies on 25 unseen\nlarge-scale Java projects. Besides assessing the correctness, we\nalso assess the diversity and strength of the generated oracles.\nWe compare the results against EvoSuite and the state-of-the-art\nneural method, TOGA. Our findings reveal that TOGLL can\nproduce 3.8 times more correct assertion oracles and 4.9 times\nmore exception oracles than TOGA. Regarding bug detection\neffectiveness, TOGLL can detect 1,023 unique mutants that\nEvoSuite cannot, which is ten times more than what TOGA can\ndetect. Additionally, TOGLL significantly outperforms TOGA in\ndetecting real bugs from the Defects4J dataset.",
          "bbox": [
            48.46403503417969,
            316.5411071777344,
            300.5245361328125,
            505.8346862792969
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 417
        },
        {
          "id": "R16-E02",
          "number": 16,
          "page": 7,
          "label": "RQ2 Finding：3.8 倍與 4.9 倍",
          "image": "assets/ref-16-e02.png",
          "source_sha256": "25c607fa39bb553e4e6a3387a49fa6a478a3087eea202534f54858d9da5b169f",
          "image_sha256": "0c7d0328542e26ae31d031bb5342383a9a0b7c077abca63e684a145662415026",
          "text": "RQ2 Findings: TOGLL generates significantly more\ncorrect test oracles than TOGA; bettering it by 3.8\ntimes and 4.9 times for assertion oracles and exception\noracles, respectively.",
          "bbox": [
            64.05400085449219,
            488.17950439453125,
            284.9314270019531,
            535.0989990234375
          ],
          "mode": "paragraph-crop",
          "width": 487,
          "height": 105
        }
      ],
      "identity_image": "assets/ref-16-identity.png"
    },
    {
      "number": 17,
      "key": "hossain2026docvscode",
      "bib": {
        "author": "Hossain, Soneya Binta and Dwyer, Matthew B. and Tasnim, Tasfia",
        "title": "Documentation vs. Code Patterns: What Drives LLM-Based Exception Oracle Generation?",
        "booktitle": "Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE)",
        "address": "Munich, Germany",
        "year": "2026",
        "doi": "10.1145/3832783.3837464"
      },
      "url": "https://doi.org/10.1145/3832783.3837464",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/hossain-2026-documentation-vs-code-oracle/auto/hossain-2026-documentation-vs-code-oracle_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "c9b761f2f1fe01d83fc23f65e2ace522915dff9f13958cf9c849d310886d43ed",
      "pages": 12,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816115117",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "resolved",
        "note": "0.54 pp 有原文依據（第 6 頁 RQ2 Finding：ΔEC −0.54 pp；整體 Δμ −0.16 pp）。v3.1 寫成「type-prediction accuracy … implying lexical shortcuts」，未交代兩個限定；v3.2 已改為「exception-versus-assertion prediction accuracy by at most 0.54 percentage points on clause-bearing samples, and further ablations reveal reliance on shortcut cues」：指標是 Exception／Assertion 二分類，數字限於含 exceptional-behavior clauses 的子集，shortcut 結論另由 attribution-guided ablation 支持（第 1 頁摘要）。整體 0.16 pp 未寫入正文。",
        "extra_links": [
          [
            "arXiv 正式紀錄與 related DOI",
            "https://arxiv.org/abs/2608.00884"
          ],
          [
            "作者研究室出版清單",
            "https://assert-lab.github.io/publications/"
          ]
        ],
        "passages": [
          {
            "page": 2,
            "blocks": [
              12,
              13
            ],
            "label": "Introduction：指標定義、整體與子集差異",
            "quote": null
          },
          {
            "page": 6,
            "blocks": [
              36
            ],
            "label": "RQ2 Finding：Δμ −0.16 pp；ΔEC −0.54 pp",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              12
            ],
            "label": "Abstract：shortcut 結論另有 attribution ablation 支持",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2608.00884v1",
      "fulltext_url": "https://arxiv.org/pdf/2608.00884v1",
      "record_url": "https://arxiv.org/abs/2608.00884v1",
      "evidence": [
        {
          "id": "R17-E01",
          "number": 17,
          "page": 2,
          "label": "Introduction：指標定義、整體與子集差異",
          "image": "assets/ref-17-e01.png",
          "source_sha256": "c9b761f2f1fe01d83fc23f65e2ace522915dff9f13958cf9c849d310886d43ed",
          "image_sha256": "b2698f99c38c1d66b70caf69fae3b3ddea3c9aeeb4620734a32953ae474475d2",
          "text": "In this paper, we investigate what actually drives exception-\noracle prediction in neural and LLM-based TOG. We answer this\nthrough a large-scale intervention-based study of three representa-\ntive TOG systems on three real-world Java datasets: two generated-\ntest benchmarks (OE25 and Sf110) and OE25dev, a newly curated\nbenchmark of developer-written tests from 25 systems. We define\nexception-oracle accuracy as an oracle-type prediction metric: for an\nexception-labeled instance, a prediction is correct when the instance\nis classified as Exception, not Assertion. This metric does not\nmeasure semantic equivalence, compilability, oracle style, or exact\nexception-type matching. Section 2.4 gives the formal definition.\nTo identify the signals behind these predictions, we compare\nmodel behavior before and after removing Javadoc @throws clauses,\nand then apply attribution-guided substitution ablations with cue\ncategorization across the test prefix, focal code, and documentation.\nRemoving these exception clauses (ECs) causes no change or only\nvery small drops in accuracy: the largest observed overall drop is\n0.16 percentage points, and the largest drop among samples that\noriginally contained exceptional-behavior clauses is 0.54 percentage\npoints. This suggests that structured exception documentation is\nnot the main driver. Instead, attribution-guided ablation shows that",
          "bbox": [
            53.07400131225586,
            480.8724365234375,
            296.05938720703125,
            710.061279296875
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 506
        },
        {
          "id": "R17-E02",
          "number": 17,
          "page": 6,
          "label": "RQ2 Finding：Δμ −0.16 pp；ΔEC −0.54 pp",
          "image": "assets/ref-17-e02.png",
          "source_sha256": "c9b761f2f1fe01d83fc23f65e2ace522915dff9f13958cf9c849d310886d43ed",
          "image_sha256": "5d449ef1ec039156f4751430692d384fe8848a9d622bb3c9a2a7878898a536a7",
          "text": "RQ2 Finding:\nJavadoc exceptional-behavior clauses provide only a weak\nauxiliary signal for exception-oracle prediction. Removing\nECs causes no change or only small drops in accuracy, with\nthe largest observed Δ𝜇being -0.16 percentage points and\nthe largest observed Δ𝐸𝐶being -0.54 percentage points for\nDoc2OracLL on OE25.",
          "bbox": [
            332.9020080566406,
            478.6202087402344,
            543.3482055664062,
            556.4797973632812
          ],
          "mode": "paragraph-crop",
          "width": 464,
          "height": 173
        },
        {
          "id": "R17-E03",
          "number": 17,
          "page": 1,
          "label": "Abstract：shortcut 結論另有 attribution ablation 支持",
          "image": "assets/ref-17-e03.png",
          "source_sha256": "c9b761f2f1fe01d83fc23f65e2ace522915dff9f13958cf9c849d310886d43ed",
          "image_sha256": "4d681a8d268332b706153c7665001d4b5cf0a3da5ce6bb58264b738aa85ddd67",
          "text": "We investigate this question through a large-scale intervention-\nbased study of three TOG systems spanning classifier-based and\ngenerative architectures and model sizes from roughly 110M to 7B\nparameters, evaluated on three real-world benchmarks comprising\ntwo generated-test datasets and a new benchmark of developer-\nwritten tests. We first remove Javadoc @throws clauses and find\nthat accuracy changes only marginally, with the largest drop be-\nlow one percentage point. This indicates that structured exception\ndocumentation is not the primary driver of exception-oracle pre-\ndiction. We then apply attribution-guided substitution ablations to\nidentify the signals that predictions depend on. The results show\nthat high accuracy can be driven by shortcut signals: some models\nare highly sensitive to a small number of structural tokens, while\nothers distribute reliance across many lexical cues.",
          "bbox": [
            52.96699905395508,
            268.32037353515625,
            296.0567626953125,
            420.7257995605469
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 336
        }
      ],
      "identity_image": "assets/ref-17-identity.png"
    },
    {
      "number": 18,
      "key": "khandaker2025augmentest",
      "bib": {
        "author": "Khandaker, Shaker Mahmud and Kifetew, Fitsum and Prandi, Davide and Susi, Angelo",
        "title": "AugmenTest: Enhancing Tests with LLM-Driven Oracles",
        "booktitle": "Proceedings of the IEEE Conference on Software Testing, Verification and Validation (ICST)",
        "address": "Napoli, Italy",
        "year": "2025",
        "doi": "10.1109/ICST62969.2025.10988926",
        "eprint": "2501.17461",
        "pages": "279--289"
      },
      "url": "https://doi.org/10.1109/ICST62969.2025.10988926",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/khandaker-2025-augmentest/auto/khandaker-2025-augmentest_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "bcbc150dec72f5be4c37ec79a884e541790f05a23bd762e5d39dc6df8db3cffc",
      "pages": 11,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916152715",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要寫明不讀 implementation code、由 documentation 與 developer comments 推論 intended behavior，並報告 Extended Prompt 30% 對 TOGA 8.2%。「documentation alone」應理解為 oracle 的語意依據，不是整個 test-generation pipeline 從不使用 test prefix 或程式工具。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              17,
              18
            ],
            "label": "Abstract：文件導向 oracle 及與 TOGA 的比較",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2501.17461v1",
      "fulltext_url": "https://arxiv.org/pdf/2501.17461v1",
      "record_url": "https://arxiv.org/abs/2501.17461v1",
      "evidence": [
        {
          "id": "R18-E01",
          "number": 18,
          "page": 1,
          "label": "Abstract：文件導向 oracle 及與 TOGA 的比較",
          "image": "assets/ref-18-e01.png",
          "source_sha256": "bcbc150dec72f5be4c37ec79a884e541790f05a23bd762e5d39dc6df8db3cffc",
          "image_sha256": "c7c4f9eed5d7bebe548b811d26775c97c9c5b3982895d664ad1354e7ce03b892",
          "text": "To address this challenge, we present AugmenTest, an ap-\nproach leveraging Large Language Models (LLMs) to infer\ncorrect test oracles based on available documentation of the\nsoftware under test. Unlike most existing methods that rely on\ncode, AugmenTest utilizes the semantic capabilities of LLMs to\ninfer the intended behavior of a method from documentation and\ndeveloper comments, without looking at the code. AugmenTest\nincludes four variants: Simple Prompt, Extended Prompt, RAG\nwith a generic prompt (without the context of class or method\nunder test), and RAG with Simple Prompt, each offering different\nlevels of contextual information to the LLMs.\nTo evaluate our work, we selected 142 Java classes and gen-\nerated multiple mutants for each. We then generated tests from\nthese mutants, focusing only on tests that passed on the mutant\nbut failed on the original class, to ensure that the tests effectively\ncaptured bugs. This resulted in 203 unique tests with distinct\nbugs, which were then used to evaluate AugmenTest. Results show\nthat in the most conservative scenario, AugmenTest’s Extended\nPrompt consistently outperformed the Simple Prompt, achieving\na success rate of 30% for generating correct assertions. In\ncomparison, the state-of-the-art TOGA approach achieved 8.2%.\nContrary to our expectations, the RAG-based approaches did not\nlead to improvements, with performance of 18.2% success rate\nfor the most conservative scenario.",
          "bbox": [
            48.46400451660156,
            269.1080322265625,
            300.5245361328125,
            508.6096496582031
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 527
        }
      ],
      "identity_image": "assets/ref-18-identity.png"
    },
    {
      "number": 19,
      "key": "ma2026reqassertions",
      "bib": {
        "author": "Ma, Tiancheng and Eisty, Nasir U.",
        "title": "From Business Requirements to Test Assertions: Evaluating LLM-Generated Oracles on Real Bugs",
        "year": "2026",
        "eprint": "2607.10277",
        "howpublished": "arXiv:2607.10277"
      },
      "url": "https://arxiv.org/abs/2607.10277",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/business_req_oracles_2607.10277/auto/business_req_oracles_2607.10277_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "ab9c22b8052a6f3b413f4d179076e316b12cf1b99a4a8c2f93d4ef74813325c0",
      "pages": 11,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816114519",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要直接記載生成 oracle 的 LLM 不取得 source code 或 input–output examples。研究者仍使用 buggy／fixed diff 建立需求，不能解讀成研究資料準備完全沒有讀程式碼。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              4,
              5,
              6
            ],
            "label": "Abstract：需求輸入、無 source code 與資料建構",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2607.10277v1",
      "fulltext_url": "https://arxiv.org/pdf/2607.10277v1",
      "record_url": "https://arxiv.org/abs/2607.10277v1",
      "evidence": [
        {
          "id": "R19-E01",
          "number": 19,
          "page": 1,
          "label": "Abstract：需求輸入、無 source code 與資料建構",
          "image": "assets/ref-19-e01.png",
          "source_sha256": "ab9c22b8052a6f3b413f4d179076e316b12cf1b99a4a8c2f93d4ef74813325c0",
          "image_sha256": "be2129a6bcfeb093aa682beba11af0e3c79188f5151583f57fe6b5948ea3d5ce",
          "text": "Background. The oracle problem (determining the correct expected outcome for a test) remains\na major bottleneck in automated testing, and is increasingly relevant as non-experts rely on AI-\ngenerated code they cannot reliably validate. Objective. We study whether large language models\n(LLMs) can generate generalizable test oracles directly from natural-language business requirements,\nwithout access to source code or example input–output pairs. Method. We propose a reproducible,\nrequirement-driven pipeline grounded in Defects4J. For each of 10 real bugs from Defects4J Lang\n(Bugs 1 and 3–11), we (i) extract behavioral changes via buggy/fixed diffs, (ii) manually translate\nthe change into a business requirement, (iii) construct a requirement-derived oracle (REQ) as a gold\nstandard, and (iv) prompt five LLMs (DeepSeek-V3, Gemma-3n, Llama-3, Mistral-7B, and Qwen-3)\nto generate Java oracle code. We evaluate oracle correctness and generalization under two targets:\nagreement with REQ and agreement with the system under test (SUT), reporting macro-averaged\naccuracy, precision, recall, and F1. Results. LLMs achieve non-trivial generalization but with\nsubstantial bug- and model-level variance. Generated oracles align more closely with REQ than with\nSUT, and correlations between requirement technicality/ambiguity ratings and oracle accuracy are\nweak with wide confidence intervals. Conclusion. No detectable linear relationship exists between\nrequirement properties and oracle accuracy in this dataset, suggesting that pretraining coverage\nand the semantic specificity of the required behavior dominate oracle correctness. As a pilot proof\nof concept, these findings are preliminary and are intended to establish feasibility and motivate\nlarger-scale empirical investigation.",
          "bbox": [
            89.12999725341797,
            209.413818359375,
            489.6334228515625,
            434.614990234375
          ],
          "mode": "paragraph-crop",
          "width": 882,
          "height": 497
        }
      ],
      "identity_image": "assets/ref-19-identity.png"
    },
    {
      "number": 20,
      "key": "wang2025mutgen",
      "bib": {
        "author": "Wang, Guancheng and Xu, Qinghua and Briand, Lionel and Liu, Kui",
        "title": "Mutation-Guided Unit Test Generation with a Large Language Model",
        "year": "2025",
        "eprint": "2506.02954",
        "howpublished": "arXiv:2506.02954"
      },
      "url": "https://arxiv.org/abs/2506.02954",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/wang-2025-mutgen/auto/wang-2025-mutgen_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "65227926805d9fea8547c24fb35077062463c0f803426fd94799b07582177172",
      "pages": 15,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816115939",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文給出 HumanEval-Java 的 id_81 實例：line／branch coverage 100%，mutation score 4%。它是存在性例子，不代表所有生成測試平均只有 4%。",
        "passages": [
          {
            "page": 2,
            "blocks": [
              45
            ],
            "label": "Running example：id_81 實例的前半段",
            "quote": null
          },
          {
            "page": 3,
            "blocks": [
              1
            ],
            "label": "Running example 續：100% coverage、4% mutation score",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2506.02954v8",
      "fulltext_url": "https://arxiv.org/pdf/2506.02954v8",
      "record_url": "https://arxiv.org/abs/2506.02954v8",
      "evidence": [
        {
          "id": "R20-E01",
          "number": 20,
          "page": 2,
          "label": "Running example：id_81 實例的前半段",
          "image": "assets/ref-20-e01.png",
          "source_sha256": "65227926805d9fea8547c24fb35077062463c0f803426fd94799b07582177172",
          "image_sha256": "c0a13d2d4193e1ef5acdf40a103fa17e12471946383d15e5f5c47bc3fe900948",
          "text": "When combined with the example code (including com-\nments), LLMs can generate test cases that achieve high line\nand branch coverage. However, as demonstrated in prior\nwork [17], [20], [21], high coverage does not necessarily\nimply strong fault-detection capability when measured by\nthe mutation score. For instance, in our experiments, LLMs\ngenerate tests for the subject id_81 from HumanEval-Java",
          "bbox": [
            311.4779968261719,
            666.096435546875,
            563.5357055664062,
            748.7900390625
          ],
          "mode": "paragraph-crop",
          "width": 555,
          "height": 183
        },
        {
          "id": "R20-E02",
          "number": 20,
          "page": 3,
          "label": "Running example 續：100% coverage、4% mutation score",
          "image": "assets/ref-20-e02.png",
          "source_sha256": "65227926805d9fea8547c24fb35077062463c0f803426fd94799b07582177172",
          "image_sha256": "4298a2e3457c924ce9839043f5012cc9b93753a5d28fc06e396eec1b97dd9ed1",
          "text": "with 100% line and branch coverage, yet the corresponding\nmutation score is only 4%.",
          "bbox": [
            48.464019775390625,
            56.38246154785156,
            300.521484375,
            79.30105590820312
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 51
        }
      ],
      "identity_image": "assets/ref-20-identity.png"
    },
    {
      "number": 21,
      "key": "chen2026agenttests",
      "bib": {
        "author": "Chen, Zhi and Sun, Zhensu and Shi, Yuling and Peng, Chao and Gu, Xiaodong and Lo, David and Jiang, Lingxiao",
        "title": "Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents",
        "year": "2026",
        "eprint": "2602.07900",
        "howpublished": "arXiv:2602.07900"
      },
      "url": "https://arxiv.org/abs/2602.07900",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/agent-generated-tests-2026/auto/agent-generated-tests-2026_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "14a983e7fb745303fd60157ee690614f8b027043aaea3c8b0c8d1b173c323cfd",
      "pages": 12,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916151707",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要直接說 agent-written tests 中 value-revealing prints 比 assertion checks 常見，研究情境為 SWE-bench Verified 的 issue resolution。本文沒有把此數字當成 Tuco 自己的結果。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              22
            ],
            "label": "Abstract：print 比 assertion checks 常見",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2602.07900v2",
      "fulltext_url": "https://arxiv.org/pdf/2602.07900v2",
      "record_url": "https://arxiv.org/abs/2602.07900v2",
      "evidence": [
        {
          "id": "R21-E01",
          "number": 21,
          "page": 1,
          "label": "Abstract：print 比 assertion checks 常見",
          "image": "assets/ref-21-e01.png",
          "source_sha256": "14a983e7fb745303fd60157ee690614f8b027043aaea3c8b0c8d1b173c323cfd",
          "image_sha256": "c199e6e3f01f294759a848e8ea3ef6fb03bc6de54e3f2ae0612e5c533930f959",
          "text": "To better understand the role of agent-written tests, we analyze\ntrajectories produced by six strong LLMs on SWE-bench Verified.\nOur results show that test writing is common, but resolved and\nunresolved tasks within the same model exhibit similar test-writing\nfrequencies. When tests are written, they mainly serve as obser-\nvational feedback channels, with value-revealing print statements\nappearing much more often than assertion-based checks. Based on\nthese insights, we perform a prompt-intervention study by revising\nthe prompts used with four models to either increase or reduce\ntest writing. The results suggest that prompt-induced changes in\nthe volume of agent-written tests do not significantly change final\noutcomes in this setting. Taken together, these results suggest that\ncurrent agent-written testing practices reshape process and cost\nmore than final task outcomes.",
          "bbox": [
            53.07400131225586,
            416.7569274902344,
            296.05609130859375,
            569.1488037109375
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 337
        }
      ],
      "identity_image": "assets/ref-21-identity.png"
    },
    {
      "number": 22,
      "key": "molinelli2025usefuloracles",
      "bib": {
        "author": "Molinelli, Davide and Di Grazia, Luca and Martin-Lopez, Alberto and Ernst, Michael D. and Pezz{\\`e}, Mauro",
        "title": "Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased Dataset",
        "booktitle": "Proceedings of the 40th IEEE/ACM International Conference on Automated Software Engineering (ASE)",
        "year": "2025",
        "doi": "10.1109/ASE63991.2025.00031",
        "pages": "278--290"
      },
      "url": "https://doi.org/10.1109/ASE63991.2025.00031",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/molinelli-2025-useful-test-oracles/auto/molinelli-2025-useful-test-oracles_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "7d3a4f4cc01488e18a8bff509dfa8531ec6ef7c1b0fa2c108cbe7b6e191ab376",
      "pages": 13,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916153432",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文報告 post-cutoff 的 13,866 個 oracle，平均 mutation score 為 LLM 43%／human 45%；同時說明公開 benchmark 可能造成 training-data leakage。後者支持 Tuco 的一般風險討論，不證明 BookStack／PrestaShop 確實在特定模型訓練集。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              15
            ],
            "label": "Abstract：post-cutoff dataset、43%／45%",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              21,
              23
            ],
            "label": "Introduction：資料洩漏風險及資料時間界線",
            "quote": null
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://doi.org/10.1109/ASE63991.2025.00031",
      "record_url": "https://doi.org/10.1109/ASE63991.2025.00031",
      "evidence": [
        {
          "id": "R22-E01",
          "number": 22,
          "page": 1,
          "label": "Abstract：post-cutoff dataset、43%／45%",
          "image": "assets/ref-22-e01.png",
          "source_sha256": "7d3a4f4cc01488e18a8bff509dfa8531ec6ef7c1b0fa2c108cbe7b6e191ab376",
          "image_sha256": "867ff2656978659003c73a508c6bcf99b3ba004f644aa86cb897eb29ac5adfea",
          "text": "Abstract—Generation of thorough test oracles is an open prob-\nlem. Popular test case generators, like EvoSuite and Randoop,\nrely on implicit, rule-based, and regression oracles that miss\nfailures that depend on the semantics of the program under test.\nFormal specifications can yield test oracles but are expensive\nto create. Large Language Models (LLMs) have the potential\nto overcome these limitations. The few studies of using LLMs to\ngenerate test oracles use modest-sized public benchmarks, such as\nDefects4J, that are likely to be included in the LLM training data,\nwhich threatens the validity of the results. This paper presents\nan empirical study of the effectiveness of LLMs in generating\ntest oracles. Our experiments use 13,866 test oracles, from 135\nJava projects, that were created after the LLMs training cut-\noff dates. Thus, our dataset is unbiased. In our experiments,\nLLMs generated oracles with average mutation score of 43% —\nsimilar to the 45% score of human-designed test oracles. Our\nresults also indicate that the test prefix and the methods called\nin the program under test provide sufficient information to\ngenerate good oracles, while additional code context does not\nbring relevant benefits. These findings provide actionable insights\ninto using LLMs for automatic testing and highlight their current\nlimitations in generating complex oracles.",
          "bbox": [
            48.46403503417969,
            266.7210693359375,
            300.5234375,
            485.9036865234375
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 483
        },
        {
          "id": "R22-E02",
          "number": 22,
          "page": 1,
          "label": "Introduction：資料洩漏風險及資料時間界線",
          "image": "assets/ref-22-e02.png",
          "source_sha256": "7d3a4f4cc01488e18a8bff509dfa8531ec6ef7c1b0fa2c108cbe7b6e191ab376",
          "image_sha256": "3cc71e059e0906e0bb5fab7654bdaee95cf3cc67a67e1365dfc7eb8d49bfdaea",
          "text": "Previous work evaluates LLMs on modest-sized public\ndatasets [18], such as Defects4J [27]. This well-known bench-\nmark is likely to have been included in LLM training data;\nsuch data leakage may inflate performance estimates [28]–\n[30]. A recent survey [31] confirms the absence of large-scale\nevaluation to assess the effectiveness of LLMs in generating\nassertions for test cases added after the LLM’s training cut-\noff. This highlights a critical gap in understanding the true\ngeneralization ability of LLMs for generating oracles [32].\nThis paper fills the gap with a large-scale empirical study of\nconcrete test oracles generated by LLMs on a dataset designed\nto avoid leakage from the training set. We extracted 13,866\noracles from 135 open-source Java projects. All of these test\ncases were created after 2024-09-01, ensuring that the test\ncode was not in the models’ training data. We evaluated\n10 LLMs from 3 families (llama, phi, and qwen), including\ngeneral-purpose, code-specific, and reasoning-enhanced vari-\nants, with 4 different prompt configurations. Our experiments\ngenerated 610,104 oracles: 13,866 oracles per LLM-and-\nprompt pair.",
          "bbox": [
            311.4780578613281,
            361.8505554199219,
            563.5357666015625,
            623.7109985351562
          ],
          "mode": "paragraph-crop",
          "width": 555,
          "height": 577
        }
      ],
      "identity_image": "assets/ref-22-identity.png"
    },
    {
      "number": 23,
      "key": "binamungu2023bdd",
      "bib": {
        "author": "Binamungu, Leonard Peter and Maro, Salome",
        "title": "Behaviour driven development: A systematic mapping study",
        "journal": "Journal of Systems and Software",
        "volume": "203",
        "pages": "111749",
        "year": "2023",
        "doi": "10.1016/j.jss.2023.111749"
      },
      "url": "https://doi.org/10.1016/j.jss.2023.111749",
      "source_path": "/home/ubuntu/tuco-lab/tmp/citation-audit/downloads/binamungu2023bdd.pdf",
      "provenance": "downloaded_primary",
      "download_url": "https://arxiv.org/pdf/2305.05567",
      "sha256": "9537ef607cf43c3a62fdf6d9badfed0afe4d662b61f5e79afb35ebf15bfdec3d",
      "pages": 65,
      "pdf_metadata": {
        "format": "PDF 1.5",
        "title": "Behaviour Driven Development: A Systematic Mapping Study",
        "author": "Leonard Peter Binamungu ; Salome Maro ; ",
        "subject": "",
        "keywords": "",
        "creator": "LaTeX with hyperref",
        "producer": "pdfTeX-1.40.21",
        "creationDate": "D:20230510004620Z",
        "modDate": "D:20230510004620Z",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "resolved",
        "note": "原文支持 BDD 以自然語言 scenarios 連接需求與可執行測試。v3.1 寫「remains the standard bridge」，帶有唯一或公認標準的語感；v3.2 已改為「links natural-language requirements to executable checks」，未新增文獻。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              6,
              7,
              8,
              9
            ],
            "label": "Abstract：自然語言規格可執行",
            "quote": "The resulting natural language specifications can also be executed to reveal correct and problematic parts of a software."
          }
        ]
      },
      "source_version": "arXiv:2305.05567v1",
      "fulltext_url": "https://arxiv.org/pdf/2305.05567v1",
      "record_url": "https://arxiv.org/abs/2305.05567v1",
      "evidence": [
        {
          "id": "R23-E01",
          "number": 23,
          "page": 1,
          "label": "Abstract：自然語言規格可執行",
          "image": "assets/ref-23-e01.png",
          "source_sha256": "9537ef607cf43c3a62fdf6d9badfed0afe4d662b61f5e79afb35ebf15bfdec3d",
          "image_sha256": "5bdf17dbd4a546fbe0e8a8e89c33786a86d24a8fa0b294ffb19227c0a0d030ee",
          "text": "The resulting natural language specifications can also be executed to reveal correct and problematic parts of a software.",
          "bbox": [
            131.76800537109375,
            336.0869445800781,
            479.52752685546875,
            367.9825439453125
          ],
          "rects": [
            [
              284.78106689453125,
              338.0869445800781,
              477.52752685546875,
              348.0495300292969
            ],
            [
              133.76800537109375,
              356.01995849609375,
              459.40545654296875,
              365.9825439453125
            ]
          ],
          "mode": "short-quote-lines",
          "width": 870,
          "height": 80
        }
      ]
    },
    {
      "number": 24,
      "key": "folorunsho2026survey",
      "bib": {
        "author": "Folorunsho, Orimoloye and Reza, Hassan",
        "title": "AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps",
        "year": "2026",
        "eprint": "2606.06563",
        "howpublished": "arXiv:2606.06563"
      },
      "url": "https://arxiv.org/abs/2606.06563",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/folorunsho-2026-ai-test-generation-survey/auto/folorunsho-2026-ai-test-generation-survey_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "61fde34b52c890187fc1c3a0a90f48a2e0c161dc042f925e2a1cda2ae99ef2a2",
      "pages": 23,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916152203",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要與研究問題對應表明列 hallucination、traceability 等 gaps。本文將它作為 survey 對研究缺口的整理，引用用途正確。",
        "passages": [
          {
            "page": 4,
            "blocks": [
              15
            ],
            "label": "研究問題對應：G1 hallucination 與 G2 traceability",
            "quote": null
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://arxiv.org/abs/2606.06563",
      "record_url": "https://arxiv.org/abs/2606.06563",
      "evidence": [
        {
          "id": "R24-E01",
          "number": 24,
          "page": 4,
          "label": "研究問題對應：G1 hallucination 與 G2 traceability",
          "image": "assets/ref-24-e01.png",
          "source_sha256": "61fde34b52c890187fc1c3a0a90f48a2e0c161dc042f925e2a1cda2ae99ef2a2",
          "image_sha256": "b2609e4efa540cba1482e51950815c8c7d948f102dda462a213df5c611180e5c",
          "text": "These research questions are operationalized by explicitly \nmapping them to the survey’s evidence base. RQ1 \n(techniques) is addressed in Sections IV and V through the \ncorpus of twenty-six primary studies organized across the \nthree evolutionary eras. RQ2 (tools and frameworks) is \naddressed in Section VI, where Table III summarizes related \nwork and provides a cross-cutting comparison of tools. RQ3 \n(evaluation) is addressed in Section VI through Table IV, \nparticularly criterion K5 on evaluation thoroughness. RQ4 \n(gaps and challenges) is addressed in Section VII through \nfour formalized research gaps: G1 (hallucination), G2 \n(traceability), \nG3 \n(complexity \nsensitivity), \nand \nG4 \n(compliance). Each gap is grounded in specific cell values in \nTable IV and in external quantitative baselines from [16] and \n[53] and is mapped one-to-one in Fig. 7 to the four actionable \nrecommendations (R1-R4) developed in Section VIII. This \nexplicit chain from research question to section, table, gap, \nand recommendation ensures that each finding remains \ntraceable to its evidentiary basis.",
          "bbox": [
            306.2200012207031,
            402.8386535644531,
            553.1438598632812,
            621.8772583007812
          ],
          "mode": "paragraph-crop",
          "width": 544,
          "height": 483
        }
      ],
      "identity_image": "assets/ref-24-identity.png"
    },
    {
      "number": 25,
      "key": "segura2016metamorphic",
      "bib": {
        "author": "Segura, Sergio and Fraser, Gordon and Sanchez, Ana B. and Ruiz-Cort{\\'e}s, Antonio",
        "title": "A Survey on Metamorphic Testing",
        "journal": "IEEE Transactions on Software Engineering",
        "volume": "42",
        "number": "9",
        "pages": "805--824",
        "year": "2016",
        "doi": "10.1109/TSE.2016.2532875"
      },
      "url": "https://doi.org/10.1109/TSE.2016.2532875",
      "source_path": "/home/ubuntu/tuco-lab/tmp/citation-audit/downloads/segura2016metamorphic.pdf",
      "provenance": "downloaded_primary",
      "download_url": "https://idus.us.es/server/api/core/bitstreams/d8fffc33-4eb5-4bb4-a3d6-c4ddab3f7cb6/content",
      "sha256": "2e7996d42db8bf5bfc163ab6ad36cc897bf3fe0376d486aadcb10e47f88b1a0b",
      "pages": 20,
      "pdf_metadata": {
        "format": "PDF 1.5",
        "title": "TSE2532875.pdf",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "Appligent AppendPDF Pro 5.5",
        "producer": "pdfTeX-1.40.13",
        "creationDate": "D:20160308132410-05'00'",
        "modDate": "D:20160308132410-05'00'",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "analogy",
        "note": "原文支持 metamorphic testing 以輸入／輸出關係檢查行為。把 Tuco 的 recorded round trip 視為弱形式的 partial oracle，是本文的設計類比；這篇 survey 沒有評估 Tuco，也不保證任意儲存往返都是有效 metamorphic relation。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              8
            ],
            "label": "Introduction：輸出關係的基本想法",
            "quote": "it is simpler to reason about relations between outputs of a program, than it is to fully understand or formalise its input-output behaviour."
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://idus.us.es/server/api/core/bitstreams/d8fffc33-4eb5-4bb4-a3d6-c4ddab3f7cb6/content",
      "record_url": "https://doi.org/10.1109/TSE.2016.2532875",
      "evidence": [
        {
          "id": "R25-E01",
          "number": 25,
          "page": 1,
          "label": "Introduction：輸出關係的基本想法",
          "image": "assets/ref-25-e01.png",
          "source_sha256": "2e7996d42db8bf5bfc163ab6ad36cc897bf3fe0376d486aadcb10e47f88b1a0b",
          "image_sha256": "ef5c09a716864ef791b641902ccabe0f2147b01e90c818c5b63c9c5475c565ec",
          "text": "it is simpler to reason about relations between outputs of a program, than it is to fully understand or formalise its input-output behaviour.",
          "bbox": [
            46.0,
            505.24322509765625,
            301.9970397949219,
            541.8231811523438
          ],
          "rects": [
            [
              48.0,
              507.24322509765625,
              299.9970397949219,
              516.7432250976562
            ],
            [
              48.0,
              518.783203125,
              299.99700927734375,
              528.283203125
            ],
            [
              48.0,
              530.3231811523438,
              151.5594940185547,
              539.8231811523438
            ]
          ],
          "mode": "short-quote-lines",
          "width": 640,
          "height": 92
        }
      ]
    },
    {
      "number": 26,
      "key": "bures2020injection",
      "bib": {
        "author": "Bures, Miroslav and Herout, Pavel and Ahmed, Bestoun S.",
        "title": "Open-source Defect Injection Benchmark Testbed for the Evaluation of Testing",
        "booktitle": "Proceedings of the IEEE International Conference on Software Testing, Verification and Validation (ICST)",
        "year": "2020",
        "pages": "442--447",
        "doi": "10.1109/ICST46399.2020.00059",
        "eprint": "2001.09342"
      },
      "url": "https://doi.org/10.1109/ICST46399.2020.00059",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/bures-2020-defect-injection-benchmark/auto/bures-2020-defect-injection-benchmark_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "fc5975987728fddb4598017af6ea848103f6884493f995d64b647fcbdf034c4c",
      "pages": 6,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916152002",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文直接支持人工缺陷植入可補足 mutation operators 的限制，也要求研究者考慮缺陷代表性。本文的 seeds 限制是把這個方法學疑慮套用到自身研究；不是原文已對 Tuco 的 30 顆缺陷做過代表性判定。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              12
            ],
            "label": "Abstract：誤解規格的複雜缺陷與 injection",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              21
            ],
            "label": "Introduction：complement，不是 replacement",
            "quote": null
          },
          {
            "page": 6,
            "blocks": [
              1,
              2
            ],
            "label": "§IV：代表性與預定缺陷集合的限制",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2001.09342v1",
      "fulltext_url": "https://arxiv.org/pdf/2001.09342v1",
      "record_url": "https://arxiv.org/abs/2001.09342v1",
      "evidence": [
        {
          "id": "R26-E01",
          "number": 26,
          "page": 1,
          "label": "Abstract：誤解規格的複雜缺陷與 injection",
          "image": "assets/ref-26-e01.png",
          "source_sha256": "fc5975987728fddb4598017af6ea848103f6884493f995d64b647fcbdf034c4c",
          "image_sha256": "ee82876075dbecc46a2487f5d19be0811953cd7f69822b2b922caa44259c6038",
          "text": "Abstract—A natural method to evaluate the effectiveness of\na testing technique is to measure the defect detection rate\nwhen applying the created test cases. Here, real or artiﬁcial\nsoftware defects can be injected into the source code of software.\nFor a more extensive evaluation, injection of artiﬁcial defects\nis usually needed and can be performed via mutation testing\nusing code mutation operators. However, to simulate complex\ndefects arising from a misunderstanding of design speciﬁcations,\nmutation testing might reach its limit in some cases. In this paper,\nwe present an open-source benchmark testbed application that\nemploys a complement method of artiﬁcial defect injection. The\napplication is compiled after artiﬁcial defects are injected into\nits source code from predeﬁned building blocks. The majority\nof the functions and user interface elements are covered by\ncreating front-end-based automated test cases that can be used\nin experiments.",
          "bbox": [
            48.46399688720703,
            218.9010772705078,
            300.5233459472656,
            378.3065490722656
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 352
        },
        {
          "id": "R26-E02",
          "number": 26,
          "page": 1,
          "label": "Introduction：complement，不是 replacement",
          "image": "assets/ref-26-e02.png",
          "source_sha256": "fc5975987728fddb4598017af6ea848103f6884493f995d64b647fcbdf034c4c",
          "image_sha256": "890b6579f5672625f2aa1d1d7eb003e404d4db81f59482ab2f59edeeb6c70492",
          "text": "In contrast to the established classical code mutation op-\nerators, various complex software defects can be introduced\ninto the code, especially defects caused by a misunderstanding\nof the SUT design speciﬁcation or requirements during the\ndevelopment process. The practical use case of the presented\ntestbed is to provide researchers with a complementary option\nto the mutation testing technique to be able to simulate a\nbroader spectrum of possible software defects during experi-\nments. The testbed is, hence, a complement to mutation testing\nrather a replacement of mutation testing via a defect injection\napproach. As we show later in Section II, both approaches have\ncertain advantages and disadvantages. Hence, both approaches\ncan be combined to provide the best objective measurement\nof the effectiveness of a testing technique.",
          "bbox": [
            311.4779968261719,
            552.9713745117188,
            563.53564453125,
            719.3519897460938
          ],
          "mode": "paragraph-crop",
          "width": 555,
          "height": 367
        },
        {
          "id": "R26-E03",
          "number": 26,
          "page": 6,
          "label": "§IV：代表性與預定缺陷集合的限制",
          "image": "assets/ref-26-e03.png",
          "source_sha256": "fc5975987728fddb4598017af6ea848103f6884493f995d64b647fcbdf034c4c",
          "image_sha256": "f4adf41bc02e4b5aebe1873bacceff196eac92cc6914974997fd7e900a62004f",
          "text": "Concern whether the introduced defects represent typical\ndefects that are being made during real software projects can\nbe raised. This responsibility in experiments is up to the re-\nsearchers and testing practitioners. Typical defects might vary\nbetween various software architectures, development styles,\nprogramming languages, business domains, and even decades\nwhen the empirical observations are made. Hence, the testbed\nprovides a general possibility to create different types of\ndefects and defect clones, and the decision is up to the testbed\nuser.\nIn the proposed concept, the artiﬁcial defects are selected\nfrom a pre-deﬁned set, which might limit the generalization\nof experiment results. This potential limit can be solved by\nthe addition of more artiﬁcial defects as well as the correct\ninterpretation of the results of the experiments.",
          "bbox": [
            48.4640007019043,
            110.73847961425781,
            300.5220031738281,
            289.18023681640625
          ],
          "mode": "paragraph-crop",
          "width": 556,
          "height": 394
        }
      ],
      "identity_image": "assets/ref-26-identity.png"
    },
    {
      "number": 27,
      "key": "gyimesi2021bugsjs",
      "bib": {
        "author": "Gyimesi, P{\\'e}ter and Vancsics, B{\\'e}la and Stocco, Andrea and Mazinanian, Davood and Besz{\\'e}des, {\\'A}rp{\\'a}d and Ferenc, Rudolf and Mesbah, Ali",
        "title": "BugsJS: A Benchmark and Taxonomy of JavaScript Bugs",
        "journal": "Software Testing, Verification and Reliability",
        "year": "2021",
        "volume": "31",
        "pages": "e1751",
        "number": "4",
        "doi": "10.1002/stvr.1751"
      },
      "url": "https://doi.org/10.1002/stvr.1751",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/gyimesi-2021-bugsjs/auto/gyimesi-2021-bugsjs_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "509e0b8313a33be0480c57e42a52a4b880ebcafe8f847478aca8bfc36734825e",
      "pages": 38,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916152413",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "analogy",
        "note": "BugsJS 原文證明有 453 個人工驗證、可重現的 server-side JavaScript bugs 及其 taxonomy。它支撐 benchmark 的先例；它本身不能證明 Tuco 的 PHP seeds 逐項來自該 taxonomy。本文的自有缺陷設計來源仍須由本研究資料說明。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              5
            ],
            "label": "Summary：453 個真實、人工驗證的 server-side bugs 與 taxonomy",
            "quote": null
          }
        ]
      },
      "source_version": "來源 PDF（以 SHA-256 固定版本）",
      "fulltext_url": "https://doi.org/10.1002/stvr.1751",
      "record_url": "https://doi.org/10.1002/stvr.1751",
      "evidence": [
        {
          "id": "R27-E01",
          "number": 27,
          "page": 1,
          "label": "Summary：453 個真實、人工驗證的 server-side bugs 與 taxonomy",
          "image": "assets/ref-27-e01.png",
          "source_sha256": "509e0b8313a33be0480c57e42a52a4b880ebcafe8f847478aca8bfc36734825e",
          "image_sha256": "20c0c1a70a9ed49424fb370b530556aeb96d0cc1322352513cb03fbcc23b8d54",
          "text": "JavaScript is a popular programming language that is also error‐prone due to its asynchronous, dynamic,\nand loosely typed nature. In recent years, numerous techniques have been proposed for analyzing and\ntesting JavaScript applications. However, our survey of the literature in this area revealed that the proposed\ntechniques are often evaluated on different datasets of programs and bugs. The lack of a commonly used\nbenchmark limits the ability to perform fair and unbiased comparisons for assessing the efﬁcacy of new\ntechniques. To ﬁll this gap, we propose BUGSJS, a benchmark of 453 real, manually validated JavaScript\nbugs from 10 popular JavaScript server‐side programs, comprising 444k lines of code (LOC) in total. Each\nbug is accompanied by its bug report, the test cases that expose it, as well as the patch that ﬁxes it. We\nextended BUGSJS with a rich web interface for visualizing and dissecting the bugs’ information, as well as\na programmable API to access the faulty and ﬁxed versions of the programs and to execute the correspond-\ning test cases, which facilitates conducting highly reproducible empirical studies and comparisons of\nJavaScript analysis and testing tools. Moreover, following a rigorous procedure, we performed a classiﬁca-\ntion of the bugs according to their nature. Our internal validation shows that our taxonomy is adequate for\ncharacterizing the bugs in BUGSJS. We discuss several ways in which the resulting taxonomy and the\nbenchmark can help direct researchers interested in automated testing of JavaScript applications. © 2021\nThe Authors. Software Testing, Veriﬁcation & Reliability published by John Wiley & Sons, Ltd.",
          "bbox": [
            93.49688720703125,
            299.40203857421875,
            500.51922607421875,
            460.2587585449219
          ],
          "mode": "paragraph-crop",
          "width": 897,
          "height": 355
        }
      ],
      "identity_image": "assets/ref-27-identity.png"
    },
    {
      "number": 28,
      "key": "tan2026agentchaos",
      "bib": {
        "author": "Tan, Gou and Sun, Zhensu and Shi, Jieke and Zhang, Ting and He, Zilong and Wu, Qingfu and Liang, Shuai and Sun, Weifeng and He, Junda and Chen, Pengfei and Zhang, Chuanfu and Shar, Lwin Khin and Lo, David",
        "title": "AgentChaos: Chaos Engineering for Agent Systems via Programmatic Fault Injection",
        "booktitle": "Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE)",
        "address": "Munich, Germany",
        "year": "2026",
        "doi": "10.1145/3832783.3837437"
      },
      "url": "https://doi.org/10.1145/3832783.3837437",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/agentchaos-2026/auto/agentchaos-2026_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "f8981089ee0278c17db1b3350a7ed7c0b2d69ad1f935bd04b57d921b43ce97a5",
      "pages": 13,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916151840",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "analogy",
        "note": "AgentChaos 的 trigger verification 有直接原文。其對象是執行期間注入的 LLM API faults，且排除當次未觸發的 task；Tuco 則在植入階段確認 GUI 缺陷可觸發，探索未找到仍保留分母。因此這是「確認注入有效」的跨情境類比，不是相同的評分或排除規則。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              36
            ],
            "label": "Abstract：LLM API fault injection 與 trigger check",
            "quote": null
          },
          {
            "page": 6,
            "blocks": [
              9,
              10,
              11
            ],
            "label": "§4.4 左欄：確認是否真正觸發",
            "quote": null
          },
          {
            "page": 6,
            "blocks": [
              12
            ],
            "label": "§4.4 右欄續：untriggered tasks 的排除規則",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2608.06790v1",
      "fulltext_url": "https://arxiv.org/pdf/2608.06790v1",
      "record_url": "https://arxiv.org/abs/2608.06790v1",
      "evidence": [
        {
          "id": "R28-E01",
          "number": 28,
          "page": 1,
          "label": "Abstract：LLM API fault injection 與 trigger check",
          "image": "assets/ref-28-e01.png",
          "source_sha256": "f8981089ee0278c17db1b3350a7ed7c0b2d69ad1f935bd04b57d921b43ce97a5",
          "image_sha256": "8962712f74385739a73d2df30f80b2a28c0d40908e458a13e86218e82750ab03",
          "text": "Agent systems rely on LLM APIs for every response, but these\nAPIs can return server errors, truncated responses, or corrupted\ncontent that propagates through downstream agents and causes\ntask failure. Evaluating robustness under these faults is crucial for\nreliable deployment. Existing fault injection methods are offline, re-\nquire source code modification, or cannot modify specific response\nfields. A comprehensive evaluation also requires a systematic fault\ntaxonomy because different fault types affect downstream agents\ndifferently. We propose AgentChaos, a chaos engineering frame-\nwork for controlled, runtime, non-intrusive LLM API fault injection.\nSince all agent systems access LLMs through the same HTTP inter-\nface, we inject faults at this shared layer without modifying source\ncode. We define crash, omission, and value faults on content and\ntool call fields, intercept and modify LLM API responses at runtime,\nand verify whether each fault is triggered to filter untriggered tasks\nand avoid underestimating fault impact. Evaluations across agent",
          "bbox": [
            52.96699905395508,
            426.7234191894531,
            296.0567932128906,
            601.1486206054688
          ],
          "mode": "paragraph-crop",
          "width": 536,
          "height": 385
        },
        {
          "id": "R28-E02",
          "number": 28,
          "page": 6,
          "label": "§4.4 左欄：確認是否真正觸發",
          "image": "assets/ref-28-e02.png",
          "source_sha256": "f8981089ee0278c17db1b3350a7ed7c0b2d69ad1f935bd04b57d921b43ce97a5",
          "image_sha256": "b64fc8e56622c232cc909e5261d71c9702e207c452eca3f1312101964a6754fb",
          "text": "4.4\nTrigger Verification\nAn agent system decides how many LLM calls to make at runtime,\nand this number varies across tasks. A fault configured for a specific\ncall position may not fire if the task finishes before reaching that\nposition. For example, if a fault is configured for the 3rd LLM call but\nthe task finishes in 2 calls, the fault is never applied. The intermittent\nstrategy may also not select any call during a short task. The target\nfield guard (§4.3) may cause additional skips if the LLM consistently\nreturns tool calls when a content fault is configured. Including such\nuntriggered tasks in the evaluation would mix faulted and unfaulted\nresults, making the system appear more robust than it actually is.\nWe therefore check after each task whether the execution trace\ncontains at least one fault event (as recorded by the wrapper in",
          "bbox": [
            52.98400115966797,
            564.3585815429688,
            295.532958984375,
            710.1060180664062
          ],
          "mode": "paragraph-crop",
          "width": 535,
          "height": 322
        },
        {
          "id": "R28-E03",
          "number": 28,
          "page": 6,
          "label": "§4.4 右欄續：untriggered tasks 的排除規則",
          "image": "assets/ref-28-e03.png",
          "source_sha256": "f8981089ee0278c17db1b3350a7ed7c0b2d69ad1f935bd04b57d921b43ce97a5",
          "image_sha256": "c80756b6ac42e1648b415b609187a066a3f5acb486334cf931d635a6673d85ec",
          "text": "§4.3). If no fault event exists, the task is marked as untriggered\nand excluded from the evaluation. Only triggered tasks are used to\ncompute pass@1w/ FI and Δpass@1. This filtering ensures that the\nmeasured degradation reflects the true impact of the injected fault.",
          "bbox": [
            317.4549865722656,
            86.35242462158203,
            560.0798950195312,
            129.2515869140625
          ],
          "mode": "paragraph-crop",
          "width": 535,
          "height": 96
        }
      ],
      "identity_image": "assets/ref-28-identity.png"
    },
    {
      "number": 29,
      "key": "zhao2026coveragemutation",
      "bib": {
        "author": "Zhao, Junda and Zhou, Shurui and Cohen, Eldan",
        "title": "Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)",
        "journal": "Proceedings of the ACM on Software Engineering",
        "volume": "3",
        "number": "ISSTA",
        "year": "2026",
        "pages": "ISSTA002",
        "doi": "10.1145/3832093"
      },
      "url": "https://doi.org/10.1145/3832093",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/zhao-2026-coverage-mutation-correlation/auto/zhao-2026-coverage-mutation-correlation_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "175c66f3c4dbf03b6034a344945324c0bb1cf50eac89e6d9c8ec6ec11515995f",
      "pages": 24,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816113729",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "§1 與 §6 直接寫出 buggy-code setting 下 coverage 不可靠、mutation 不適用；§3 註腳解釋標準 mutation analysis 要求 green test suite，會排除本來就揭露 bug 的 tests。本文已有「once code already contains faults」限定；不能脫離該設定概括所有 mutation 方法。",
        "passages": [
          {
            "page": 3,
            "blocks": [
              1
            ],
            "label": "Introduction：buggy-code setting 的結論",
            "quote": null
          },
          {
            "page": 5,
            "blocks": [
              12
            ],
            "label": "§3 註腳：為何標準 mutation analysis 不適用",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2607.22880v1",
      "fulltext_url": "https://arxiv.org/pdf/2607.22880v1",
      "record_url": "https://arxiv.org/abs/2607.22880v1",
      "evidence": [
        {
          "id": "R29-E01",
          "number": 29,
          "page": 3,
          "label": "Introduction：buggy-code setting 的結論",
          "image": "assets/ref-29-e01.png",
          "source_sha256": "175c66f3c4dbf03b6034a344945324c0bb1cf50eac89e6d9c8ec6ec11515995f",
          "image_sha256": "aa6cf642903cfa4e462870549d0b41e82f6e708a56cdbc6e8fa8eae5edac118e",
          "text": "provide useful signals when comparing across models. However, in the more general and practically\nchallenging setting where the correctness of the code-under-test provided to the LLM cannot be\nguaranteed, and the generated tests are expected to detect bugs present in that code, coverage\nbecomes unreliable as an indicator of bug detection effectiveness, and mutation analysis is not\napplicable. Furthermore, unlike both prior studies, we find little evidence that the size of LLM-\ngenerated test suites (i.e., the number of tests) is a strong confounding factor in the relationships\namong coverage, mutation score, and real-bug detection effectiveness.",
          "bbox": [
            45.327999114990234,
            86.74288940429688,
            442.35577392578125,
            169.36129760742188
          ],
          "mode": "paragraph-crop",
          "width": 875,
          "height": 183
        },
        {
          "id": "R29-E02",
          "number": 29,
          "page": 5,
          "label": "§3 註腳：為何標準 mutation analysis 不適用",
          "image": "assets/ref-29-e02.png",
          "source_sha256": "175c66f3c4dbf03b6034a344945324c0bb1cf50eac89e6d9c8ec6ec11515995f",
          "image_sha256": "838ea3069ec3ee9c995289159e2f40c4969de7672c58680ce91ad6707c45c906",
          "text": "1We do not conduct mutation testing, or analyze the correlation between mutation score and bug detection, for buggy code\nin RQ2 because mutation testing presupposes a passing (“green”) test suite on the code-under-test; this would exclude\nprecisely the tests that expose the bug within the code-under-test, hiding the very bug to be detected and rendering the\nresulting mutation scores meaningless. Mutation analysis is therefore not applicable in this setting.",
          "bbox": [
            45.191001892089844,
            618.2855834960938,
            440.67364501953125,
            658.541259765625
          ],
          "mode": "paragraph-crop",
          "width": 871,
          "height": 89
        }
      ],
      "identity_image": "assets/ref-29-identity.png"
    },
    {
      "number": 30,
      "key": "liu2026testexplora",
      "bib": {
        "author": "Liu, Steven and Luo, Jane and Zhang, Xin and Liu, Aofan and Liu, Hao and Wu, Jie and Huang, Ziyang and Huang, Yangyu and Kang, Yu and Li, Scarlett",
        "title": "TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation",
        "year": "2026",
        "eprint": "2602.10471",
        "howpublished": "arXiv:2602.10471"
      },
      "url": "https://arxiv.org/abs/2602.10471",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/testexplora-2026/auto/testexplora-2026_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "d8ab766eaf6fa5e750c22b23bb2412d6a05db3e4a7ca4ebd4abe364688e29e64",
      "pages": 43,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916154012",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "arXiv v2 摘要直接說 hides all defect-related signals，§5.1 明列所有 baseline experiments 重複三次。這是作者對 benchmark 設計的描述；本文不是主張 independently 證明毫無資訊洩漏。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              6
            ],
            "label": "Abstract：proactive discovery 與隱藏 defect signals",
            "quote": null
          },
          {
            "page": 6,
            "blocks": [
              4
            ],
            "label": "§5.1：baseline experiments 重複三次",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2602.10471v2",
      "fulltext_url": "https://arxiv.org/pdf/2602.10471v2",
      "record_url": "https://arxiv.org/abs/2602.10471v2",
      "evidence": [
        {
          "id": "R30-E01",
          "number": 30,
          "page": 1,
          "label": "Abstract：proactive discovery 與隱藏 defect signals",
          "image": "assets/ref-30-e01.png",
          "source_sha256": "d8ab766eaf6fa5e750c22b23bb2412d6a05db3e4a7ca4ebd4abe364688e29e64",
          "image_sha256": "641512494678bc9f12acb383557598c6394fa3d473611c45e75c516112d8ac11",
          "text": "Given that Large Language Models (LLMs) are\nincreasingly applied to automate software devel-\nopment, comprehensive software assurance spans\nthree distinct goals: regression prevention, reac-\ntive reproduction, and proactive discovery. Cur-\nrent evaluations systematically overlook the third\ngoal. Specifically, they either treat existing code\nas ground truth (a compliance trap) for regres-\nsion prevention, or depend on post-failure arti-\nfacts (e.g., issue reports) for bug reproduction—so\nthey rarely surface defects before failures. To\nbridge this gap, we present TestExplora, a bench-\nmark designed to evaluate LLMs as proactive\ntesters within full-scale, realistic repository envi-\nronments. TestExplora contains 2,389 tasks from\n482 repositories and hides all defect-related sig-\nnals. Models must proactively find bugs by com-\nparing implementations against documentation-\nderived intent, using documentation as the ora-\ncle. Furthermore, to keep evaluation sustainable\nand reduce leakage, we propose continuous, time-\naware data collection. Our evaluation reveals a\nsignificant capability gap: state-of-the-art models\nachieve a maximum Fail-to-Pass (F2P) rate of\nonly 16.06%. Further analysis indicates that nav-\nigating complex cross-module interactions and\nleveraging agentic exploration are critical to ad-\nvancing LLMs toward autonomous software qual-\nity assurance. Consistent with this, SWEAgent\ninstantiated with GPT-5-mini achieves an F2P of\n17.27% and an F2P@5 of 29.7%, highlighting\nthe effectiveness and promise of agentic explo-\nration in proactive bug discovery tasks.",
          "bbox": [
            74.11799621582031,
            274.7073059082031,
            271.67315673828125,
            668.31103515625
          ],
          "mode": "paragraph-crop",
          "width": 435,
          "height": 867
        },
        {
          "id": "R30-E02",
          "number": 30,
          "page": 6,
          "label": "§5.1：baseline experiments 重複三次",
          "image": "assets/ref-30-e02.png",
          "source_sha256": "d8ab766eaf6fa5e750c22b23bb2412d6a05db3e4a7ca4ebd4abe364688e29e64",
          "image_sha256": "94baf73a62cee57de93bff1ea86675ab2ffa3ef1e261c80c2b7da7f0d66a7e02",
          "text": "We conduct a comprehensive evaluation utilizing four key\nmetrics: Head Pass Rate, Fail-to-Pass Rate, Entry Coverage,\nand Change-focused Coverage. Additionally, the number\nof testcases (Num.) is reported. Our experiments involve\nsix representative language models evaluated on 12,227\nreal-world pull requests. To ensure evaluation efficiency,\nwe constructed a high-quality subset, TestExplora-Lite, by\nfiltering samples based on the quality of human-written\ndocstrings. This subset comprises 330 PRs and 517 samples\nin total. All baseline experiments are repeated three times\nto ensure the robustness of the results.",
          "bbox": [
            54.47200012207031,
            169.0422821044922,
            291.1879577636719,
            299.6160888671875
          ],
          "mode": "paragraph-crop",
          "width": 522,
          "height": 289
        }
      ],
      "identity_image": "assets/ref-30-identity.png"
    },
    {
      "number": 31,
      "key": "norman2026reliabilitywithoutvalidity",
      "bib": {
        "author": "Norman, Justin D. and Rivera, Michael U. and Hughes, D. Alex",
        "title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
        "year": "2026",
        "eprint": "2606.19544",
        "howpublished": "arXiv:2606.19544"
      },
      "url": "https://arxiv.org/abs/2606.19544",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/norman-2026-llm-judge-reliability/auto/norman-2026-llm-judge-reliability_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "aef5580c01fb71653a416892d855c04bbc4f69e5b49039dda22f3cb5491a0a84",
      "pages": 25,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260816115720",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "摘要同時支持 raw agreement 未校正 chance，以及高 test–retest reliability 可與 position bias 共存；跨 benchmark 排名差異支撐 task dependence。不能用此研究推算 Tuco grader 的實際錯誤率。",
        "passages": [
          {
            "page": 1,
            "blocks": [
              4
            ],
            "label": "Abstract：chance correction、consistency–bias paradox",
            "quote": null
          },
          {
            "page": 1,
            "blocks": [
              11
            ],
            "label": "Introduction：跨任務變化與系統性偏差",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2606.19544v1",
      "fulltext_url": "https://arxiv.org/pdf/2606.19544v1",
      "record_url": "https://arxiv.org/abs/2606.19544v1",
      "evidence": [
        {
          "id": "R31-E01",
          "number": 31,
          "page": 1,
          "label": "Abstract：chance correction、consistency–bias paradox",
          "image": "assets/ref-31-e01.png",
          "source_sha256": "aef5580c01fb71653a416892d855c04bbc4f69e5b49039dda22f3cb5491a0a84",
          "image_sha256": "1d5cd5b9c904a2787cc692f1cc4e790885c9903e4d3663965a22dadf244501c7",
          "text": "LLM-as-a-Judge has become the dominant\nevaluation paradigm for language models, but\njudge validation in practice relies on exact-\nmatch agreement, a metric that does not cor-\nrect for chance and systematically overstates\ndiscriminative ability. We present the largest\nsystematic evaluation of LLM-as-a-Judge to\ndate: 21 judges from nine providers across MT-\nBench, JudgeBench, and RewardBench, eval-\nuated under three protocols (agreement, con-\nsistency, bias audit) over 118 runs and approx-\nimately 541,000 individual judgments. Four\nfindings emerge, consistent across the full co-\nhort, including the April 2026 frontier: kappa\ndeflation between exact match and Cohen’s κ is\nuniversal (33–41 pp on MT-Bench), judge rank-\nings shift by up to 14 positions across bench-\nmarks, high test–retest reliability (> 0.95) co-\nexists with severe position bias (> 0.10) in\ntwo production-deployed judges (instantiating\na consistency–bias paradox), and verbosity bias\nis small (< 0.011) across our cohort under a\nsingle pairwise rubric. We distill these into a\nMinimum Viable Validation Protocol.",
          "bbox": [
            87.3740005493164,
            243.28628540039062,
            274.2830810546875,
            529.2930297851562
          ],
          "mode": "paragraph-crop",
          "width": 412,
          "height": 630
        },
        {
          "id": "R31-E02",
          "number": 31,
          "page": 1,
          "label": "Introduction：跨任務變化與系統性偏差",
          "image": "assets/ref-31-e02.png",
          "source_sha256": "aef5580c01fb71653a416892d855c04bbc4f69e5b49039dda22f3cb5491a0a84",
          "image_sha256": "d1851c69084ef3368b878d1eaf4a41568468d8f6c01099822cd0f5a2d4809475",
          "text": "We report five principal findings, including two\ndiagnostic concepts. First, kappa deflation: raw\nagreement overstates chance-corrected discrimina-\ntion by 33–41pp in all 21 evaluated models. Sec-\nond, judge rankings are non-transferable: models\nshift by as many as 14 positions across benchmarks.\nThird, the consistency–bias paradox: high test–\nretest reliability often masks severe position bias.\nFourth, verbosity bias is much reduced: all 21 mod-\nels register <0.011, in sharp contrast to the 20–\n40% variance reported in 2023 literature. Fifth,\nJudgeBench discriminates 4.5× more sharply than\nMT-Bench (60.4pp vs. 13.5pp κ spread).",
          "bbox": [
            305.260009765625,
            533.7122802734375,
            526.8213500976562,
            708.2237548828125
          ],
          "mode": "paragraph-crop",
          "width": 489,
          "height": 385
        }
      ],
      "identity_image": "assets/ref-31-identity.png"
    },
    {
      "number": 32,
      "key": "han2025judgesverdict",
      "bib": {
        "author": "Han, Steve and Titericz Junior, Gilberto and Balough, Tom and Zhou, Wenfei",
        "title": "Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement",
        "year": "2025",
        "eprint": "2510.09738",
        "howpublished": "arXiv:2510.09738"
      },
      "url": "https://arxiv.org/abs/2510.09738",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/judges-verdict-2025/auto/judges-verdict-2025_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "06a641d3a5b90a1ed59883d736fdfa5554f2a45f040dc1e7d5f26a0406a1a017",
      "pages": 15,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916151448",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "supported",
        "note": "原文直接說 correlation alone 不足，並舉出高度相關但持續過嚴／過寬的 judge。§4.7 的 task／prompt variation 是與 [31,33] 合引的一般理由；本篇最直接支持的是 correlation 與 validity 的區分。",
        "passages": [
          {
            "page": 2,
            "blocks": [
              1
            ],
            "label": "Introduction：correlation 不等於 agreement／validity",
            "quote": null
          },
          {
            "page": 2,
            "blocks": [
              3
            ],
            "label": "Method：不同 prompt 順序與 positional bias",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2510.09738v1",
      "fulltext_url": "https://arxiv.org/pdf/2510.09738v1",
      "record_url": "https://arxiv.org/abs/2510.09738v1",
      "evidence": [
        {
          "id": "R32-E01",
          "number": 32,
          "page": 2,
          "label": "Introduction：correlation 不等於 agreement／validity",
          "image": "assets/ref-32-e01.png",
          "source_sha256": "06a641d3a5b90a1ed59883d736fdfa5554f2a45f040dc1e7d5f26a0406a1a017",
          "image_sha256": "b8ec6c3d984ff92ca9fb8c02b16b7bcdc9662f4d78b383825020a26a34a2b932",
          "text": "Pearson’s r) to evaluate judge quality, we demonstrate that correlation alone is insufficient. Our\nmethodology progresses to Cohen’s Kappa, which measures actual agreement rather than just linear\nrelationships. This addresses critical issues like systematic bias—an LLM could have perfect cor-\nrelation while consistently being too harsh or lenient. Second, we design a novel Turing Test for\njudges based on Cohen’s Kappa agreement patterns. Unlike traditional Turing Tests that focus on\nconversational indistinguishability, our approach asks: “When mixed with human annotators, can\nwe distinguish the LLM from typical human judges?” This test uses z-score analysis of Cohen’s\nKappa values to identify models that judge like typical human annotators (|z| < 1) versus those\nwith exceptional consistency patterns. These innovations establish a more rigorous framework for\nvalidating LLM judgment excellence, moving beyond superficial correlation to actual functional\ncapability, forming a new benchmark—the Judge’s Verdict Benchmark—that provides a standard-\nized way to assess whether an LLM achieves Tier 1 performance for either human-like evaluation\nor maximum-consistency tasks.",
          "bbox": [
            107.5,
            83.53788757324219,
            504.50335693359375,
            226.23788452148438
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 315
        },
        {
          "id": "R32-E02",
          "number": 32,
          "page": 2,
          "label": "Method：不同 prompt 順序與 positional bias",
          "image": "assets/ref-32-e02.png",
          "source_sha256": "06a641d3a5b90a1ed59883d736fdfa5554f2a45f040dc1e7d5f26a0406a1a017",
          "image_sha256": "98c316cb8c1c0e054f093426803428938f9a9e4ce63c59842af596903a575ab6",
          "text": "To assess the alignment of each LLM-as-a-judge, we used the Answer Accuracy metric from the\nRAGAS1 library which measures how closely a generated answer matches a reference answer. It\nemploys a Large Language Model (LLM) as an evaluator, or “judge,” to assess the factual and\nsemantic concordance between the reference and the generated response. The process involves\ntwo independent LLM-as-a-judge prompts, each of which are prompted with the user’s question,\nthe system-generated answer, and the reference answer in different orders. Each judge assigns a\ndiscrete score of: 0 (No Alignment), 2 (Partial Alignment), or 4 (Exact Alignment) 1. To derive\nthe final score, these discrete ratings are first normalized to a continuous scale in 2. The normalized\nscores of the two judges are then averaged to produce a final accuracy score of the generated answer\nin 3. This diverse method improves the reliability of the evaluation by mitigating the positional bias\nand increases the robustness of a single LLM-as-a-judge prompt. More formally, let Si ∈{0, 2, 4}\ndenote the discrete score of the LLM judge i comparing the reference and the generated response,\nwhere:",
          "bbox": [
            107.5,
            261.644287109375,
            504.50439453125,
            404.1137390136719
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 315
        }
      ],
      "identity_image": "assets/ref-32-identity.png"
    },
    {
      "number": 33,
      "key": "dev2026judgeharness",
      "bib": {
        "author": "Dev, Sunishchal and Sloan, Andrew and Kavner, Joshua and Kong, Nicholas and Sandler, Morgan",
        "title": "Judge Reliability Harness: Stress Testing the Reliability of LLM Judges",
        "year": "2026",
        "eprint": "2603.05399v1",
        "howpublished": "arXiv:2603.05399v1",
        "url": "https://arxiv.org/abs/2603.05399v1"
      },
      "url": "https://arxiv.org/abs/2603.05399v1",
      "source_path": "/home/ubuntu/mypaper2/papers-extracted/judge-reliability-harness-2026/auto/judge-reliability-harness-2026_origin.pdf",
      "provenance": "user_library",
      "download_url": "",
      "sha256": "75e474f8e49a8e91e94a3e31db65af8cbcf26fa38bc0ec891b540cdfe8507f7d",
      "pages": 13,
      "pdf_metadata": {
        "format": "PDF 1.7",
        "title": "",
        "author": "",
        "subject": "",
        "keywords": "",
        "creator": "PDFium",
        "producer": "PDFium",
        "creationDate": "D:20260916152618",
        "modDate": "",
        "trapped": "",
        "encryption": null
      },
      "review": {
        "status": "resolved",
        "note": "v3.1 書目寫成 ICLR 主會議 proceedings；官方 arXiv 紀錄 Comments 欄明列接受於 Agents in the Wild: Safety, Security, and Beyond Workshop at ICLR 2026。v3.2 改為 @misc 預印本書目，明列 arXiv:2603.05399v1 與 URL，保留該版 LLM Judges 題名；本文三處引用內容在 v1 原文均可對到（摘要及 §6），截圖與頁碼沿用。未改引 OpenReview 的 workshop 版本（題名為 AI Judges），若日後改引須重新對齊題名、會名與頁碼。",
        "metadata_note": "v3.2 書目：Dev et al., Judge Reliability Harness: Stress Testing the Reliability of LLM Judges, arXiv:2603.05399v1, 2026, https://arxiv.org/abs/2603.05399v1。v3.1 的 ICLR proceedings 記載已移除。",
        "extra_links": [
          [
            "arXiv 紀錄（Comments 列出 workshop）",
            "https://arxiv.org/abs/2603.05399"
          ],
          [
            "正式 workshop 版本",
            "https://openreview.net/pdf?id=hajH5OBXF8"
          ]
        ],
        "passages": [
          {
            "page": 1,
            "blocks": [
              4
            ],
            "label": "arXiv v1 Abstract：無 uniformly reliable judge",
            "quote": null
          },
          {
            "page": 8,
            "blocks": [
              2,
              3
            ],
            "label": "arXiv v1 §6：task dependence 與 formatting perturbation",
            "quote": null
          }
        ]
      },
      "source_version": "arXiv:2603.05399v1",
      "fulltext_url": "https://arxiv.org/pdf/2603.05399v1",
      "record_url": "https://arxiv.org/abs/2603.05399v1",
      "evidence": [
        {
          "id": "R33-E01",
          "number": 33,
          "page": 1,
          "label": "arXiv v1 Abstract：無 uniformly reliable judge",
          "image": "assets/ref-33-e01.png",
          "source_sha256": "75e474f8e49a8e91e94a3e31db65af8cbcf26fa38bc0ec891b540cdfe8507f7d",
          "image_sha256": "52128e95790c5fff71b83d4566c796bd2f46faf6caeb469e1c82079a40f11e1d",
          "text": "We present the Judge Reliability Harness, an open source library for constructing\nvalidation suites that test the reliability of LLM judges. As LLM based scoring is\nwidely deployed in AI benchmarks, more tooling is needed to efficiently assess the\nreliability of these methods. Given a benchmark dataset and an LLM judge config-\nuration, the harness generates reliability tests that evaluate both binary judgment\naccuracy and ordinal grading performance for free-response and agentic task for-\nmats. We evaluate four state-of-the-art judges across four benchmarks spanning\nsafety, persuasion, misuse, and agentic behavior, and find meaningful variation\nin performance across models and perturbation types, highlighting opportunities\nto improve the robustness of LLM judges. No judge that we evaluated is uni-\nformly reliable across benchmarks using our harness. For example, our prelimi-\nnary experiments on judges revealed consistency issues as measured by accuracy\nin judging another LLM’s ability to complete a task due to simple text format-\nting changes, paraphrasing, changes in verbosity, and flipping the ground truth\nlabel in LLM-produced responses. The code for this tool is available at: https:\n//github.com/RANDCorporation/judge-reliability-harness",
          "bbox": [
            143.364990234375,
            241.38746643066406,
            468.6376647949219,
            416.22283935546875
          ],
          "mode": "paragraph-crop",
          "width": 717,
          "height": 385
        },
        {
          "id": "R33-E02",
          "number": 33,
          "page": 8,
          "label": "arXiv v1 §6：task dependence 與 formatting perturbation",
          "image": "assets/ref-33-e02.png",
          "source_sha256": "75e474f8e49a8e91e94a3e31db65af8cbcf26fa38bc0ec891b540cdfe8507f7d",
          "image_sha256": "ca65439e49170bf5144c390c692349b55ae01ae703cbdfdcd95a83506d2f011b",
          "text": "Judge output robustness is highly task-dependent. Models that appear stable in binary safety-\nclassification settings (e.g., FORTRESS or HarmBench) degrade substantially when required to\nassign multi-level ordinal scores, as in Persuade.\nPractitioners who rely on ordinal scoring or\npreference-ranking tasks may be overestimating the reliability of their evaluation systems.\nFormatting perturbations produce larger reliability drops than semantic perturbations. This\nasymmetry is concerning, as different LLMs tend to have unique quirks in how they format their\nresponses. Judges that are brittle to such differences risk embedding instability into downstream\nmodel comparisons or leaderboard decisions, even when semantic content remains unchanged.",
          "bbox": [
            107.5,
            111.57347869873047,
            504.5034484863281,
            205.31692504882812
          ],
          "mode": "paragraph-crop",
          "width": 874,
          "height": 207
        }
      ],
      "identity_image": "assets/ref-33-identity.png"
    }
  ],
  "occurrences": [
    {
      "id": "C01",
      "page": 1,
      "numbers": [
        1
      ],
      "marker": "[1]",
      "context_before": "ction favoured the disciplined arm in three batches and reversed in the fourth, and no\n          seeded defect was confirmed by the disciplined arm alone.\n\n         Keywords:  large language models; exploratory web testing; browser agents; test oracles;\n          evidence contract; fault injection; seeded defects; LLM graders\n\n                           1.  INTRODUCTION\n\n     Exploratory testing of a web application asks a tester to operate the product, decide\n what each screen ought to have shown, and report what did not match. Language-model\n agents now do the operating: they drive real browsers over self-hosted open-source ap-\n plications ",
      "context_after": " and over production-grade projects [2], and they return fluent prose about\n what they did. Two gaps remain at the ends of that chain.\n    The first i",
      "file": "01-introduction.tex",
      "line": 5,
      "keys": [
        "zhou2024webarena"
      ],
      "section": "§1 Introduction",
      "quote": "they drive real browsers over self-hosted open-source applications",
      "evidence_links": {
        "1": [
          "R01-E01",
          "R01-E02"
        ]
      }
    },
    {
      "id": "C02",
      "page": 1,
      "numbers": [
        2
      ],
      "marker": "[2]",
      "context_before": "hree batches and reversed in the fourth, and no\n          seeded defect was confirmed by the disciplined arm alone.\n\n         Keywords:  large language models; exploratory web testing; browser agents; test oracles;\n          evidence contract; fault injection; seeded defects; LLM graders\n\n                           1.  INTRODUCTION\n\n     Exploratory testing of a web application asks a tester to operate the product, decide\n what each screen ought to have shown, and report what did not match. Language-model\n agents now do the operating: they drive real browsers over self-hosted open-source ap-\n plications [1] and over production-grade projects ",
      "context_after": ", and they return fluent prose about\n what they did. Two gaps remain at the ends of that chain.\n    The first is the expectation gap. Oracles derived ",
      "file": "01-introduction.tex",
      "line": 6,
      "keys": [
        "teoh2026webtestpilot"
      ],
      "section": "§1 Introduction",
      "quote": "over production-grade projects",
      "evidence_links": {
        "2": [
          "R02-E02"
        ]
      }
    },
    {
      "id": "C03",
      "page": 1,
      "numbers": [
        3,
        4
      ],
      "marker": "[3, 4]",
      "context_before": "ded defects; LLM graders\n\n                           1.  INTRODUCTION\n\n     Exploratory testing of a web application asks a tester to operate the product, decide\n what each screen ought to have shown, and report what did not match. Language-model\n agents now do the operating: they drive real browsers over self-hosted open-source ap-\n plications [1] and over production-grade projects [2], and they return fluent prose about\n what they did. Two gaps remain at the ends of that chain.\n    The first is the expectation gap. Oracles derived from code tend to encode the be-\n haviour a program exhibits rather than the behaviour a specification demands ",
      "context_after": ", and\n over half of the oracle studies in a recent review issue verdicts with no grounding in a spec-\n ification artefact [5]. An agent that reports a",
      "file": "01-introduction.tex",
      "line": 12,
      "keys": [
        "konstantinou2024actualexpected",
        "bodicoat2025understanding"
      ],
      "section": "§1 Introduction",
      "quote": "Oracles derived from code tend to encode the behaviour a program exhibits rather than the behaviour a specification demands",
      "evidence_links": {
        "3": [
          "R03-E01"
        ],
        "4": [
          "R04-E01"
        ]
      }
    },
    {
      "id": "C04",
      "page": 1,
      "numbers": [
        5
      ],
      "marker": "[5]",
      "context_before": "er to operate the product, decide\n what each screen ought to have shown, and report what did not match. Language-model\n agents now do the operating: they drive real browsers over self-hosted open-source ap-\n plications [1] and over production-grade projects [2], and they return fluent prose about\n what they did. Two gaps remain at the ends of that chain.\n    The first is the expectation gap. Oracles derived from code tend to encode the be-\n haviour a program exhibits rather than the behaviour a specification demands [3, 4], and\n over half of the oracle studies in a recent review issue verdicts with no grounding in a spec-\n ification artefact ",
      "context_after": ". An agent that reports a surprise without stating what it expected, and\n on what basis, leaves a reader nothing to check the surprise against.\n\n +Cor",
      "file": "01-introduction.tex",
      "line": 15,
      "keys": [
        "mughal2026sourceofauthority"
      ],
      "section": "§1 Introduction",
      "quote": "over half of the oracle studies in a recent review issue verdicts with no grounding in a specification artefact",
      "evidence_links": {
        "5": [
          "R05-E01"
        ]
      }
    },
    {
      "id": "C05",
      "page": 2,
      "numbers": [
        6
      ],
      "marker": "[6]",
      "context_before": "2                            C.-S. KOONG, Y.-C. LIN, Y.-T. CHEN\n\n\n\n    The second is the evidence gap. A report keeps the symptom and loses the trajectory:\nthe reader receives a paragraph asserting that an export failed, without the operations\nand captures that would show it, and an agent optimizing for completion also walks past\nanomalies it did not need ",
      "context_after": ". A claim no evidence can confirm costs a reviewer as much\ntime as the defect it reports.\n    Tuco was proposed as an end-to-end answer to both gaps. ",
      "file": "01-introduction.tex",
      "line": 23,
      "keys": [
        "gao2026guitester"
      ],
      "section": "§1 Introduction",
      "quote": "an agent optimizing for completion also walks past anomalies it did not need",
      "evidence_links": {
        "6": [
          "R06-E01"
        ]
      }
    },
    {
      "id": "C06",
      "page": 2,
      "numbers": [
        7
      ],
      "marker": "[7]",
      "context_before": "he disciplined\n     arm over 128 graded candidates against 80.0% for the baseline arm over 120, or\n     88.1% once eleven environment-class candidates are removed. Detection favours\n      the disciplined arm in three batches and reverses in the fourth, no seeded defect\n     was confirmed by the disciplined arm alone, and the cost comparison is bounded\n     by runs whose cost record is incomplete.\n    This article extends the Tuco prototype introduced at TCSE, focusing on its princi-\nples of reviewable evidence and justified judgments. Building on the conference paper’s\npreliminary demonstration of a requirement-to-acceptance-testing pipeline ",
      "context_after": ", we study\na recorded-exploration workflow derived from those principles. The extension adds a con-\ntrolled evaluation over two applications, 30 seede",
      "file": "01-introduction.tex",
      "line": 65,
      "keys": [
        "koong2026tuco"
      ],
      "section": "§1 Introduction",
      "quote": "the conference paper’s preliminary demonstration of a requirement-to-acceptance-testing pipeline",
      "evidence_links": {
        "7": [
          "R07-E01"
        ]
      }
    },
    {
      "id": "C07",
      "page": 2,
      "numbers": [
        1
      ],
      "marker": "[1]",
      "context_before": "ng on the conference paper’s\npreliminary demonstration of a requirement-to-acceptance-testing pipeline [7], we study\na recorded-exploration workflow derived from those principles. The extension adds a con-\ntrolled evaluation over two applications, 30 seeded defects and four experimental batches,\nexamining candidate precision, defect detection and execution cost. The new experiments\nevaluate this derived workflow.\n\n                          2. RELATED WORK\n\n2.1 LLM Agents for Web and GUI Testing\n\n    WebArena introduced a self-hostable environment built from open-source applica-\ntions for browser agents driven by natural-language instructions ",
      "context_after": ".  Later work asks",
      "file": "02-related-work.tex",
      "line": 6,
      "keys": [
        "zhou2024webarena"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "WebArena introduced a self-hostable environment built from open-source applications for browser agents driven by natural-language instructions",
      "evidence_links": {
        "1": [
          "R01-E01",
          "R01-E02"
        ]
      }
    },
    {
      "id": "C08",
      "page": 3,
      "numbers": [
        8
      ],
      "marker": "[8]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases ",
      "context_after": ", NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI ",
      "file": "02-related-work.tex",
      "line": 8,
      "keys": [
        "chevrot2025pinata"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "PinATA quantifies automation errors and hallucinated step validations on manual test cases",
      "evidence_links": {
        "8": [
          "R08-E01",
          "R08-E02"
        ]
      }
    },
    {
      "id": "C09",
      "page": 3,
      "numbers": [
        9
      ],
      "marker": "[9]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering ",
      "context_after": ", Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constraine",
      "file": "02-related-work.tex",
      "line": 10,
      "keys": [
        "shahbandeh2024naviqate"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "NaviQAte reframes exploration as question answering",
      "evidence_links": {
        "9": [
          "R09-E01"
        ]
      }
    },
    {
      "id": "C10",
      "page": 3,
      "numbers": [
        10
      ],
      "marker": "[10]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers ",
      "context_after": " miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form m",
      "file": "02-related-work.tex",
      "line": 11,
      "keys": [
        "mesbah2012crawljax"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "Temac reaches functionality classic crawlers [10] miss [11]",
      "evidence_links": {
        "10": [
          "R10-E01"
        ]
      }
    },
    {
      "id": "C11",
      "page": 3,
      "numbers": [
        11
      ],
      "marker": "[11]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss ",
      "context_after": ", WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form model reaso",
      "file": "02-related-work.tex",
      "line": 11,
      "keys": [
        "liu2025temac"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "Temac reaches functionality classic crawlers [10] miss [11]",
      "evidence_links": {
        "11": [
          "R11-E01",
          "R11-E02"
        ]
      }
    },
    {
      "id": "C12",
      "page": 3,
      "numbers": [
        2
      ],
      "marker": "[2]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form model reasoning ",
      "context_after": ", and requirement-oriented\nvariants verify requirements on a running mobile application [12] or on a GUI prototype\n[13]. Every model evaluated under W",
      "file": "02-related-work.tex",
      "line": 14,
      "keys": [
        "teoh2026webtestpilot"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "WebTestPilot symbolizes GUI elements so that assertions are generated in a constrained DSL over those symbols rather than by free-form model reasoning",
      "evidence_links": {
        "2": [
          "R02-E01"
        ]
      }
    },
    {
      "id": "C13",
      "page": 3,
      "numbers": [
        12
      ],
      "marker": "[12]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form model reasoning [2], and requirement-oriented\nvariants verify requirements on a running mobile application ",
      "context_after": " or on a GUI prototype\n[13]. Every model evaluated under WebTestBench’s harness scores below 30% end-to-\nend F1; most sit near 30% precision against r",
      "file": "02-related-work.tex",
      "line": 16,
      "keys": [
        "hu2024auitestagent"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "requirement-oriented variants verify requirements on a running mobile application",
      "evidence_links": {
        "12": [
          "R12-E01"
        ]
      }
    },
    {
      "id": "C14",
      "page": 3,
      "numbers": [
        13
      ],
      "marker": "[13]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING             3\n\n\n\nwhether such an agent can test: PinATA quantifies automation errors and hallucinated\nstep validations on manual test cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form model reasoning [2], and requirement-oriented\nvariants verify requirements on a running mobile application [12] or on a GUI prototype\n",
      "context_after": ". Every model evaluated under WebTestBench’s harness scores below 30% end-to-\nend F1; most sit near 30% precision against recall under 25%, though the",
      "file": "02-related-work.tex",
      "line": 17,
      "keys": [
        "kolthoff2025guispector"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "or on a GUI prototype",
      "evidence_links": {
        "13": [
          "R13-E01"
        ]
      }
    },
    {
      "id": "C15",
      "page": 3,
      "numbers": [
        14
      ],
      "marker": "[14]",
      "context_before": "st cases [8], NaviQAte reframes exploration as question\nanswering [9], Temac reaches functionality classic crawlers [10] miss [11], WebTestPi-\nlot symbolizes GUI elements so that assertions are generated in a constrained DSL over\nthose symbols rather than by free-form model reasoning [2], and requirement-oriented\nvariants verify requirements on a running mobile application [12] or on a GUI prototype\n[13]. Every model evaluated under WebTestBench’s harness scores below 30% end-to-\nend F1; most sit near 30% precision against recall under 25%, though the strongest trades\nprecision for recall, on a protocol and denominators that differ from ours ",
      "context_after": ". GUITester\nnames goal-oriented masking, an agent suppressing anomalies that did not block com-\npletion, and Execution-Bias Attribution, a product def",
      "file": "02-related-work.tex",
      "line": 21,
      "keys": [
        "kong2026webtestbench"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "Every model evaluated under WebTestBench’s harness scores below 30% end-to-end F1; most sit near 30% precision against recall under 25%, though the strongest trades precision for recall, on a protocol and denominators that differ from ours",
      "evidence_links": {
        "14": [
          "R14-E01",
          "R14-E02"
        ]
      }
    },
    {
      "id": "C16",
      "page": 3,
      "numbers": [
        6
      ],
      "marker": "[6]",
      "context_before": "are generated in a constrained DSL over\nthose symbols rather than by free-form model reasoning [2], and requirement-oriented\nvariants verify requirements on a running mobile application [12] or on a GUI prototype\n[13]. Every model evaluated under WebTestBench’s harness scores below 30% end-to-\nend F1; most sit near 30% precision against recall under 25%, though the strongest trades\nprecision for recall, on a protocol and denominators that differ from ours [14]. GUITester\nnames goal-oriented masking, an agent suppressing anomalies that did not block com-\npletion, and Execution-Bias Attribution, a product defect blamed on the agent’s mis-click\n",
      "context_after": ". These systems vary the agent; we hold the worker, the browser client and the recorder\nfixed and vary the workflow package the agent follows, so a di",
      "file": "02-related-work.tex",
      "line": 24,
      "keys": [
        "gao2026guitester"
      ],
      "section": "§2.1 LLM Agents",
      "quote": "GUITester names goal-oriented masking, an agent suppressing anomalies that did not block completion, and Execution-Bias Attribution, a product defect blamed on the agent’s mis-click",
      "evidence_links": {
        "6": [
          "R06-E01"
        ]
      }
    },
    {
      "id": "C17",
      "page": 3,
      "numbers": [
        15
      ],
      "marker": "[15]",
      "context_before": "-\nend F1; most sit near 30% precision against recall under 25%, though the strongest trades\nprecision for recall, on a protocol and denominators that differ from ours [14]. GUITester\nnames goal-oriented masking, an agent suppressing anomalies that did not block com-\npletion, and Execution-Bias Attribution, a product defect blamed on the agent’s mis-click\n[6]. These systems vary the agent; we hold the worker, the browser client and the recorder\nfixed and vary the workflow package the agent follows, so a difference is not attributable\nto a stronger tester.\n\n2.2  Oracles and Expectation Grounding\n\n    The oracle problem predates language models ",
      "context_after": "; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural",
      "file": "02-related-work.tex",
      "line": 32,
      "keys": [
        "barr2015oracle"
      ],
      "section": "§2.2 Oracles",
      "quote": "The oracle problem predates language models",
      "evidence_links": {
        "15": [
          "R15-E01"
        ]
      }
    },
    {
      "id": "C18",
      "page": 3,
      "numbers": [
        16
      ],
      "marker": "[16]",
      "context_before": "ours [14]. GUITester\nnames goal-oriented masking, an agent suppressing anomalies that did not block com-\npletion, and Execution-Bias Attribution, a product defect blamed on the agent’s mis-click\n[6]. These systems vary the agent; we hold the worker, the browser client and the recorder\nfixed and vary the workflow package the agent follows, so a difference is not attributable\nto a stronger tester.\n\n2.2  Oracles and Expectation Grounding\n\n    The oracle problem predates language models [15]; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural method ",
      "context_after": ". What they encode is the worry. They capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the ",
      "file": "02-related-work.tex",
      "line": 35,
      "keys": [
        "hossain2025togll"
      ],
      "section": "§2.2 Oracles",
      "quote": "TOGLL producing 3.8 times more correct assertion oracles than TOGA, the prior state-of-the-art neural method",
      "evidence_links": {
        "16": [
          "R16-E01",
          "R16-E02"
        ]
      }
    },
    {
      "id": "C19",
      "page": 3,
      "numbers": [
        3
      ],
      "marker": "[3]",
      "context_before": "d not block com-\npletion, and Execution-Bias Attribution, a product defect blamed on the agent’s mis-click\n[6]. These systems vary the agent; we hold the worker, the browser client and the recorder\nfixed and vary the workflow package the agent follows, so a difference is not attributable\nto a stronger tester.\n\n2.2  Oracles and Expectation Grounding\n\n    The oracle problem predates language models [15]; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural method [16]. What they encode is the worry. They capture\nactual rather than expected behaviour ",
      "context_after": "; prompting and supplied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertio",
      "file": "02-related-work.tex",
      "line": 37,
      "keys": [
        "konstantinou2024actualexpected"
      ],
      "section": "§2.2 Oracles",
      "quote": "They capture actual rather than expected behaviour",
      "evidence_links": {
        "3": [
          "R03-E01"
        ]
      }
    },
    {
      "id": "C20",
      "page": 3,
      "numbers": [
        4
      ],
      "marker": "[4]",
      "context_before": "on the agent’s mis-click\n[6]. These systems vary the agent; we hold the worker, the browser client and the recorder\nfixed and vary the workflow package the agent follows, so a difference is not attributable\nto a stronger tester.\n\n2.2  Oracles and Expectation Grounding\n\n    The oracle problem predates language models [15]; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural method [16]. What they encode is the worry. They capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the choice of model ",
      "context_after": "; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples,",
      "file": "02-related-work.tex",
      "line": 39,
      "keys": [
        "bodicoat2025understanding"
      ],
      "section": "§2.2 Oracles",
      "quote": "prompting and supplied context move accuracy more than the choice of model",
      "evidence_links": {
        "4": [
          "R04-E02",
          "R04-E03"
        ]
      }
    },
    {
      "id": "C21",
      "page": 3,
      "numbers": [
        17
      ],
      "marker": "[17]",
      "context_before": " a stronger tester.\n\n2.2  Oracles and Expectation Grounding\n\n    The oracle problem predates language models [15]; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural method [16]. What they encode is the worry. They capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues ",
      "context_after": "; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from busines",
      "file": "02-related-work.tex",
      "line": 42,
      "keys": [
        "hossain2026docvscode"
      ],
      "section": "§2.2 Oracles",
      "quote": "removing Javadoc @throws clauses changes exception-versus-assertion prediction accuracy by at most 0.54 percentage points on clause-bearing samples, and further ablations reveal reliance on shortcut cues",
      "evidence_links": {
        "17": [
          "R17-E01",
          "R17-E02",
          "R17-E03"
        ]
      }
    },
    {
      "id": "C22",
      "page": 3,
      "numbers": [
        18
      ],
      "marker": "[18]",
      "context_before": "ls [15]; fine-tuned code models now gen-\nerate oracles, TOGLL producing 3.8 times more correct assertion oracles than TOGA, the\nprior state-of-the-art neural method [16]. What they encode is the worry. They capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n",
      "context_after": ", and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch ",
      "file": "02-related-work.tex",
      "line": 44,
      "keys": [
        "khandaker2025augmentest"
      ],
      "section": "§2.2 Oracles",
      "quote": "assertions derived from documentation alone still beat a neural generator that reads the code",
      "evidence_links": {
        "18": [
          "R18-E01"
        ]
      }
    },
    {
      "id": "C23",
      "page": 3,
      "numbers": [
        19
      ],
      "marker": "[19]",
      "context_before": "correct assertion oracles than TOGA, the\nprior state-of-the-art neural method [16]. What they encode is the worry. They capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n",
      "context_after": ". The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where prin",
      "file": "02-related-work.tex",
      "line": 45,
      "keys": [
        "ma2026reqassertions"
      ],
      "section": "§2.2 Oracles",
      "quote": "assertions can be derived from business requirements without the source code",
      "evidence_links": {
        "19": [
          "R19-E01"
        ]
      }
    },
    {
      "id": "C24",
      "page": 3,
      "numbers": [
        20
      ],
      "marker": "[20]",
      "context_before": "hey capture\nactual rather than expected behaviour [3]; prompting and supplied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent ",
      "context_after": ", in agent trajectories where print statements outnumber\nassertions [21], and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnea",
      "file": "02-related-work.tex",
      "line": 47,
      "keys": [
        "wang2025mutgen"
      ],
      "section": "§2.2 Oracles",
      "quote": "suites with full line and branch coverage but a mutation score of a few percent",
      "evidence_links": {
        "20": [
          "R20-E01",
          "R20-E02"
        ]
      }
    },
    {
      "id": "C25",
      "page": 3,
      "numbers": [
        21
      ],
      "marker": "[21]",
      "context_before": "plied context move accu-\nracy more than the choice of model [4]; removing Javadoc @throws clauses changes\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where print statements outnumber\nassertions ",
      "context_after": ", and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnearly the same rate [22]. A review of 83 oracle studies finds over half is",
      "file": "02-related-work.tex",
      "line": 49,
      "keys": [
        "chen2026agenttests"
      ],
      "section": "§2.2 Oracles",
      "quote": "agent trajectories where print statements outnumber assertions",
      "evidence_links": {
        "21": [
          "R21-E01"
        ]
      }
    },
    {
      "id": "C26",
      "page": 3,
      "numbers": [
        22
      ],
      "marker": "[22]",
      "context_before": "anges\nexception-versus-assertion prediction accuracy by at most 0.54 percentage points on\nclause-bearing samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where print statements outnumber\nassertions [21], and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnearly the same rate ",
      "context_after": ". A review of 83 oracle studies finds over half issue verdicts with\nno specification grounding [5], and behaviour-driven development links natural-lan",
      "file": "02-related-work.tex",
      "line": 51,
      "keys": [
        "molinelli2025usefuloracles"
      ],
      "section": "§2.2 Oracles",
      "quote": "on a post-cutoff dataset where LLM and human oracles kill mutants at nearly the same rate",
      "evidence_links": {
        "22": [
          "R22-E01"
        ]
      }
    },
    {
      "id": "C27",
      "page": 3,
      "numbers": [
        5
      ],
      "marker": "[5]",
      "context_before": "aring samples, and further ablations reveal reliance on shortcut cues [17]; yet as-\nsertions derived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where print statements outnumber\nassertions [21], and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnearly the same rate [22]. A review of 83 oracle studies finds over half issue verdicts with\nno specification grounding ",
      "context_after": ", and behaviour-driven development links natural-language\nrequirements to executable checks [23], though a review of requirements-to-test-case gen-\ner",
      "file": "02-related-work.tex",
      "line": 53,
      "keys": [
        "mughal2026sourceofauthority"
      ],
      "section": "§2.2 Oracles",
      "quote": "A review of 83 oracle studies finds over half issue verdicts with no specification grounding",
      "evidence_links": {
        "5": [
          "R05-E01"
        ]
      }
    },
    {
      "id": "C28",
      "page": 3,
      "numbers": [
        23
      ],
      "marker": "[23]",
      "context_before": "rived from documentation alone still beat a neural generator that reads the code\n[18], and assertions can be derived from business requirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where print statements outnumber\nassertions [21], and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnearly the same rate [22]. A review of 83 oracle studies finds over half issue verdicts with\nno specification grounding [5], and behaviour-driven development links natural-language\nrequirements to executable checks ",
      "context_after": ", though a review of requirements-to-test-case gen-\neration still lists traceability and hallucination control as open problems [24]. The work-\nflow s",
      "file": "02-related-work.tex",
      "line": 55,
      "keys": [
        "binamungu2023bdd"
      ],
      "section": "§2.2 Oracles",
      "quote": "behaviour-driven development links natural-language requirements to executable checks",
      "evidence_links": {
        "23": [
          "R23-E01"
        ]
      }
    },
    {
      "id": "C29",
      "page": 3,
      "numbers": [
        24
      ],
      "marker": "[24]",
      "context_before": "equirements without the source code\n[19]. The same weakness recurs in suites with full line and branch coverage but a muta-\ntion score of a few percent [20], in agent trajectories where print statements outnumber\nassertions [21], and on a post-cutoff dataset where LLM and human oracles kill mutants at\nnearly the same rate [22]. A review of 83 oracle studies finds over half issue verdicts with\nno specification grounding [5], and behaviour-driven development links natural-language\nrequirements to executable checks [23], though a review of requirements-to-test-case gen-\neration still lists traceability and hallucination control as open problems ",
      "context_after": ". The work-\nflow studied here works one step upstream of a fixed authority:  it requires the agent to\nname which of three bases supports each expected",
      "file": "02-related-work.tex",
      "line": 57,
      "keys": [
        "folorunsho2026survey"
      ],
      "section": "§2.2 Oracles",
      "quote": "a review of requirements-to-test-case generation still lists traceability and hallucination control as open problems",
      "evidence_links": {
        "24": [
          "R24-E01"
        ]
      }
    },
    {
      "id": "C30",
      "page": 3,
      "numbers": [
        25
      ],
      "marker": "[25]",
      "context_before": "finds over half issue verdicts with\nno specification grounding [5], and behaviour-driven development links natural-language\nrequirements to executable checks [23], though a review of requirements-to-test-case gen-\neration still lists traceability and hallucination control as open problems [24]. The work-\nflow studied here works one step upstream of a fixed authority:  it requires the agent to\nname which of three bases supports each expected outcome and to file an unsupported\nexpectation as a gap rather than a defect. One of those bases, a recorded round trip, is a\nweak form of the partial oracles that state a property relating two executions ",
      "context_after": ".\n\n2.3  Evidence, Seeded-Defect Benchmarks and LLM Graders\n\n    Defect injection predates the language-model era. Bures et al.  built a testbed for\nit",
      "file": "02-related-work.tex",
      "line": 63,
      "keys": [
        "segura2016metamorphic"
      ],
      "section": "§2.2 Oracles",
      "quote": "One of those bases, a recorded round trip, is a weak form of the partial oracles that state a property relating two executions",
      "evidence_links": {
        "25": [
          "R25-E01"
        ]
      }
    },
    {
      "id": "C31",
      "page": 3,
      "numbers": [
        26
      ],
      "marker": "[26]",
      "context_before": "p upstream of a fixed authority:  it requires the agent to\nname which of three bases supports each expected outcome and to file an unsupported\nexpectation as a gap rather than a defect. One of those bases, a recorded round trip, is a\nweak form of the partial oracles that state a property relating two executions [25].\n\n2.3  Evidence, Seeded-Defect Benchmarks and LLM Graders\n\n    Defect injection predates the language-model era. Bures et al.  built a testbed for\nit, argued that mutation testing might reach its limit for faults that arise from a misread\nspecification, and offered injection as a complement to mutation rather than a replace-\nment ",
      "context_after": "; that is the argument for seeding semantic server-side defects, and their caveat\napplies to ours, since seeded defects cannot be shown to represent t",
      "file": "02-related-work.tex",
      "line": 72,
      "keys": [
        "bures2020injection"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "Bures et al. built a testbed for it, argued that mutation testing might reach its limit for faults that arise from a misread specification, and offered injection as a complement to mutation rather than a replacement",
      "evidence_links": {
        "26": [
          "R26-E01",
          "R26-E02",
          "R26-E03"
        ]
      }
    },
    {
      "id": "C32",
      "page": 3,
      "numbers": [
        27
      ],
      "marker": "[27]",
      "context_before": "roperty relating two executions [25].\n\n2.3  Evidence, Seeded-Defect Benchmarks and LLM Graders\n\n    Defect injection predates the language-model era. Bures et al.  built a testbed for\nit, argued that mutation testing might reach its limit for faults that arise from a misread\nspecification, and offered injection as a complement to mutation rather than a replace-\nment [26]; that is the argument for seeding semantic server-side defects, and their caveat\napplies to ours, since seeded defects cannot be shown to represent those real development\nproduces.  Validated collections of reproducible server-side bugs supply the taxonomy\nsuch seeds imitate ",
      "context_after": ", and agent evaluations inject one defect per requirement across",
      "file": "02-related-work.tex",
      "line": 76,
      "keys": [
        "gyimesi2021bugsjs"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "Validated collections of reproducible server-side bugs supply the taxonomy such seeds imitate",
      "evidence_links": {
        "27": [
          "R27-E01"
        ]
      }
    },
    {
      "id": "C33",
      "page": 4,
      "numbers": [
        2
      ],
      "marker": "[2]",
      "context_before": "4                            C.-S. KOONG, Y.-C. LIN, Y.-T. CHEN\n\n\n\nproduction-grade applications ",
      "context_after": ", label a large test-item pool against applications whose\ndefects arise from AI generation rather than injection, 1,750 items of which 448 fail [14],\n",
      "file": "02-related-work.tex",
      "line": 78,
      "keys": [
        "teoh2026webtestpilot"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "agent evaluations inject one defect per requirement across production-grade applications",
      "evidence_links": {
        "2": [
          "R02-E03"
        ]
      }
    },
    {
      "id": "C34",
      "page": 4,
      "numbers": [
        14
      ],
      "marker": "[14]",
      "context_before": "4                            C.-S. KOONG, Y.-C. LIN, Y.-T. CHEN\n\n\n\nproduction-grade applications [2], label a large test-item pool against applications whose\ndefects arise from AI generation rather than injection, 1,750 items of which 448 fail ",
      "context_after": ",\nor pair interactive tasks with a defect-type catalogue [6]. Fault-injection studies verify\nthat a seeded fault is activatable, so that an injection ",
      "file": "02-related-work.tex",
      "line": 80,
      "keys": [
        "kong2026webtestbench"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "label a large test-item pool against applications whose defects arise from AI generation rather than injection, 1,750 items of which 448 fail",
      "evidence_links": {
        "14": [
          "R14-E03",
          "R14-E04",
          "R14-E05"
        ]
      }
    },
    {
      "id": "C35",
      "page": 4,
      "numbers": [
        6
      ],
      "marker": "[6]",
      "context_before": "4                            C.-S. KOONG, Y.-C. LIN, Y.-T. CHEN\n\n\n\nproduction-grade applications [2], label a large test-item pool against applications whose\ndefects arise from AI generation rather than injection, 1,750 items of which 448 fail [14],\nor pair interactive tasks with a defect-type catalogue ",
      "context_after": ". Fault-injection studies verify\nthat a seeded fault is activatable, so that an injection with no observable effect does not\ndilute the measured effec",
      "file": "02-related-work.tex",
      "line": 82,
      "keys": [
        "gao2026guitester"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "pair interactive tasks with a defect-type catalogue",
      "evidence_links": {
        "6": [
          "R06-E02",
          "R06-E03"
        ]
      }
    },
    {
      "id": "C36",
      "page": 4,
      "numbers": [
        28
      ],
      "marker": "[28]",
      "context_before": "4                            C.-S. KOONG, Y.-C. LIN, Y.-T. CHEN\n\n\n\nproduction-grade applications [2], label a large test-item pool against applications whose\ndefects arise from AI generation rather than injection, 1,750 items of which 448 fail [14],\nor pair interactive tasks with a defect-type catalogue [6]. Fault-injection studies verify\nthat a seeded fault is activatable, so that an injection with no observable effect does not\ndilute the measured effect ",
      "context_after": ". That check concerns the seeding rather than the search: a\nseeded defect an agent never found still counts in the denominator of fifteen per system.\n",
      "file": "02-related-work.tex",
      "line": 84,
      "keys": [
        "tan2026agentchaos"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "Fault-injection studies verify that a seeded fault is activatable, so that an injection with no observable effect does not dilute the measured effect",
      "evidence_links": {
        "28": [
          "R28-E01",
          "R28-E02",
          "R28-E03"
        ]
      }
    },
    {
      "id": "C37",
      "page": 4,
      "numbers": [
        29
      ],
      "marker": "[29]",
      "context_before": "t applications whose\ndefects arise from AI generation rather than injection, 1,750 items of which 448 fail [14],\nor pair interactive tasks with a defect-type catalogue [6]. Fault-injection studies verify\nthat a seeded fault is activatable, so that an injection with no observable effect does not\ndilute the measured effect [28]. That check concerns the seeding rather than the search: a\nseeded defect an agent never found still counts in the denominator of fifteen per system.\nThe familiar proxies can mislead here: coverage loses predictive power, and mutation\nanalysis does not apply at all, once the code already contains the faults to be exposed ",
      "context_after": ",\nand TestExplora hides every defect-related signal from the agent and repeats its baseline\nconfigurations three times [30]. Because our confirmed-def",
      "file": "02-related-work.tex",
      "line": 90,
      "keys": [
        "zhao2026coveragemutation"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "coverage loses predictive power, and mutation analysis does not apply at all, once the code already contains the faults to be exposed",
      "evidence_links": {
        "29": [
          "R29-E01",
          "R29-E02"
        ]
      }
    },
    {
      "id": "C38",
      "page": 4,
      "numbers": [
        30
      ],
      "marker": "[30]",
      "context_before": "teractive tasks with a defect-type catalogue [6]. Fault-injection studies verify\nthat a seeded fault is activatable, so that an injection with no observable effect does not\ndilute the measured effect [28]. That check concerns the seeding rather than the search: a\nseeded defect an agent never found still counts in the denominator of fifteen per system.\nThe familiar proxies can mislead here: coverage loses predictive power, and mutation\nanalysis does not apply at all, once the code already contains the faults to be exposed [29],\nand TestExplora hides every defect-related signal from the agent and repeats its baseline\nconfigurations three times ",
      "context_after": ". Because our confirmed-defect counts come from a fresh\nlanguage-model session reading frozen bundles, the grader is part of the apparatus, and re-\nli",
      "file": "02-related-work.tex",
      "line": 92,
      "keys": [
        "liu2026testexplora"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "TestExplora hides every defect-related signal from the agent and repeats its baseline configurations three times",
      "evidence_links": {
        "30": [
          "R30-E01",
          "R30-E02"
        ]
      }
    },
    {
      "id": "C39",
      "page": 4,
      "numbers": [
        31
      ],
      "marker": "[31]",
      "context_before": "counts in the denominator of fifteen per system.\nThe familiar proxies can mislead here: coverage loses predictive power, and mutation\nanalysis does not apply at all, once the code already contains the faults to be exposed [29],\nand TestExplora hides every defect-related signal from the agent and repeats its baseline\nconfigurations three times [30]. Because our confirmed-defect counts come from a fresh\nlanguage-model session reading frozen bundles, the grader is part of the apparatus, and re-\nliability work bounds what such counts carry: agreement rates are uncorrected for chance\nand a grader can be consistent while carrying a systematic bias ",
      "context_after": ", correlation alone\ndoes not validate a grader [32], and under formatting and paraphrase perturbations no\nevaluated grader is uniformly reliable [33].",
      "file": "02-related-work.tex",
      "line": 97,
      "keys": [
        "norman2026reliabilitywithoutvalidity"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "agreement rates are uncorrected for chance and a grader can be consistent while carrying a systematic bias",
      "evidence_links": {
        "31": [
          "R31-E01",
          "R31-E02"
        ]
      }
    },
    {
      "id": "C40",
      "page": 4,
      "numbers": [
        32
      ],
      "marker": "[32]",
      "context_before": "e familiar proxies can mislead here: coverage loses predictive power, and mutation\nanalysis does not apply at all, once the code already contains the faults to be exposed [29],\nand TestExplora hides every defect-related signal from the agent and repeats its baseline\nconfigurations three times [30]. Because our confirmed-defect counts come from a fresh\nlanguage-model session reading frozen bundles, the grader is part of the apparatus, and re-\nliability work bounds what such counts carry: agreement rates are uncorrected for chance\nand a grader can be consistent while carrying a systematic bias [31], correlation alone\ndoes not validate a grader ",
      "context_after": ", and under formatting and paraphrase perturbations no\nevaluated grader is uniformly reliable [33]. We therefore treat the grader as a measuring\ninstr",
      "file": "02-related-work.tex",
      "line": 99,
      "keys": [
        "han2025judgesverdict"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "correlation alone does not validate a grader",
      "evidence_links": {
        "32": [
          "R32-E01"
        ]
      }
    },
    {
      "id": "C41",
      "page": 4,
      "numbers": [
        33
      ],
      "marker": "[33]",
      "context_before": "ot apply at all, once the code already contains the faults to be exposed [29],\nand TestExplora hides every defect-related signal from the agent and repeats its baseline\nconfigurations three times [30]. Because our confirmed-defect counts come from a fresh\nlanguage-model session reading frozen bundles, the grader is part of the apparatus, and re-\nliability work bounds what such counts carry: agreement rates are uncorrected for chance\nand a grader can be consistent while carrying a systematic bias [31], correlation alone\ndoes not validate a grader [32], and under formatting and paraphrase perturbations no\nevaluated grader is uniformly reliable ",
      "context_after": ". We therefore treat the grader as a measuring\ninstrument: Sec. 4 states its protocol and Sec. 7 what that protocol leaves unresolved.\n\n 3. THE EVIDEN",
      "file": "02-related-work.tex",
      "line": 101,
      "keys": [
        "dev2026judgeharness"
      ],
      "section": "§2.3 Evidence and Benchmarks",
      "quote": "under formatting and paraphrase perturbations no evaluated grader is uniformly reliable",
      "evidence_links": {
        "33": [
          "R33-E01"
        ]
      }
    },
    {
      "id": "C42",
      "page": 4,
      "numbers": [
        7
      ],
      "marker": "[7]",
      "context_before": "is part of the apparatus, and re-\nliability work bounds what such counts carry: agreement rates are uncorrected for chance\nand a grader can be consistent while carrying a systematic bias [31], correlation alone\ndoes not validate a grader [32], and under formatting and paraphrase perturbations no\nevaluated grader is uniformly reliable [33]. We therefore treat the grader as a measuring\ninstrument: Sec. 4 states its protocol and Sec. 7 what that protocol leaves unresolved.\n\n 3. THE EVIDENCE-DISCIPLINED EXPLORATORY TESTING\n                WORKFLOW\n\n3.1  Background: The Tuco Pipeline\n\n    Tuco was introduced as a prototype pipeline of four stages ",
      "context_after": ". A rule extractor\nnormalizes the supplied requirement documents into acceptance rules, each naming a\ncondition and an outcome a browser can observe, ",
      "file": "03-framework.tex",
      "line": 4,
      "keys": [
        "koong2026tuco"
      ],
      "section": "§3.1 Background",
      "quote": "Tuco was introduced as a prototype pipeline of four stages",
      "evidence_links": {
        "7": [
          "R07-E01",
          "R07-E02",
          "R07-E03"
        ]
      }
    },
    {
      "id": "C43",
      "page": 7,
      "numbers": [
        2
      ],
      "marker": "[2]",
      "context_before": "exploration and verification instructions together with a form-state checkpoint tool.\nThe state-checkpoint tool produced 55 form diffs over the disciplined arm’s sixteen runs;\ntwo reported a change in a field the operation had not targeted, four reported a field added\nafter the save, and no diff was itself scored.\n\n4.3  Subject Systems and Seeded Defects\n\n    The two subject systems are BookStack, a documentation wiki, and PrestaShop, an\ne-commerce storefront with a back office. Both are open source, publish end-user doc-\numentation that we supply unchanged as the product documentation, and are the subject\nsystems of WebTestPilot’s benchmark ",
      "context_after": "; only the subjects are shared, since the versions\nand the seeded defects differ. Excluding tests and vendored libraries, BookStack 26.05.4\nhas about ",
      "file": "04-design.tex",
      "line": 54,
      "keys": [
        "teoh2026webtestpilot"
      ],
      "section": "§4.3 Subject Systems",
      "quote": "are the subject systems of WebTestPilot’s benchmark",
      "evidence_links": {
        "2": [
          "R02-E02"
        ]
      }
    },
    {
      "id": "C44",
      "page": 7,
      "numbers": [
        26,
        27
      ],
      "marker": "[26, 27]",
      "context_before": "nd-user doc-\numentation that we supply unchanged as the product documentation, and are the subject\nsystems of WebTestPilot’s benchmark [2]; only the subjects are shared, since the versions\nand the seeded defects differ. Excluding tests and vendored libraries, BookStack 26.05.4\nhas about 70 thousand lines of first-party code and PrestaShop 9.1.5 about 524 thousand,\nspanning a small-team product and a production platform.\n    Each system carries fifteen server-side seeded defects, injected in application code\nbehind the HTTP interface, in the tradition of controlled defect injection and of validated\ncollections of reproducible server-side bugs ",
      "context_after": ". Both arms observe a defect only\nthrough the interface a user sees. The authors designed the defects, using a large language\nmodel as a brainstorming",
      "file": "04-design.tex",
      "line": 64,
      "keys": [
        "bures2020injection",
        "gyimesi2021bugsjs"
      ],
      "section": "§4.3 Subject Systems",
      "quote": "in the tradition of controlled defect injection and of validated collections of reproducible server-side bugs",
      "evidence_links": {
        "26": [
          "R26-E01",
          "R26-E02"
        ],
        "27": [
          "R27-E01"
        ]
      }
    },
    {
      "id": "C45",
      "page": 9,
      "numbers": [
        31,
        32,
        33
      ],
      "marker": "[31, 32, 33]",
      "context_before": "all clock.\n    Each cell holds two runs, so we report direction and magnitude and run no inferential\nstatistics.\n\n4.7  Grading Protocol\n\n   A fresh Claude Opus 5 grading session evaluated the shuffled run bundles of a batch\nusing the corresponding system answer keys. The grader saw the bundles and the answer\nkeys and nothing else: no method label, cost record or instruction text. It labels each can-\ndidate admitted to grading hit, maybe, miss, duplicate or ineligible, admitting a hit only if\nit names one seeded defect and cites an evidence file inside the bundle that shows the fail-\nure. Because grader reliability varies with task and prompt ",
      "context_after": ", we duplicate the\ninstrument. The reported results use one grading round for Batch 1 and two independent\ngrading rounds for each of Batches 2 to 4, w",
      "file": "04-design.tex",
      "line": 173,
      "keys": [
        "norman2026reliabilitywithoutvalidity",
        "han2025judgesverdict",
        "dev2026judgeharness"
      ],
      "section": "§4.7 Grading",
      "quote": "Because grader reliability varies with task and prompt",
      "evidence_links": {
        "31": [
          "R31-E01",
          "R31-E02"
        ],
        "32": [
          "R32-E02"
        ],
        "33": [
          "R33-E01",
          "R33-E02"
        ]
      }
    },
    {
      "id": "C46",
      "page": 15,
      "numbers": [
        31,
        32
      ],
      "marker": "[31, 32]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING            15\n\n\n\nconfiguration.\n    Grading.  Candidate judgments were produced by an LLM grader and have not\nbeen independently validated by human reviewers. Agreement between grader sessions\nis not validity: a grader can be consistent and still carry a systematic bias ",
      "context_after": ", and\nreliability degrades most under formatting perturbations, to which graders are less robust\nthan to semantic paraphrase [33]. We admitted a hit o",
      "file": "07-threats.tex",
      "line": 29,
      "keys": [
        "norman2026reliabilitywithoutvalidity",
        "han2025judgesverdict"
      ],
      "section": "§7 Threats: Grading",
      "quote": "Agreement between grader sessions is not validity: a grader can be consistent and still carry a systematic bias",
      "evidence_links": {
        "31": [
          "R31-E01",
          "R31-E02"
        ],
        "32": [
          "R32-E01"
        ]
      }
    },
    {
      "id": "C47",
      "page": 15,
      "numbers": [
        33
      ],
      "marker": "[33]",
      "context_before": "            TUCO: EVIDENCE DISCIPLINE FOR EXPLORATORY WEB TESTING            15\n\n\n\nconfiguration.\n    Grading.  Candidate judgments were produced by an LLM grader and have not\nbeen independently validated by human reviewers. Agreement between grader sessions\nis not validity: a grader can be consistent and still carry a systematic bias [31, 32], and\nreliability degrades most under formatting perturbations, to which graders are less robust\nthan to semantic paraphrase ",
      "context_after": ". We admitted a hit only when it named one seeded defect\nand cited bundle evidence, and duplicated the grader in Batches 2 to 4.\n     Partial blinding",
      "file": "07-threats.tex",
      "line": 32,
      "keys": [
        "dev2026judgeharness"
      ],
      "section": "§7 Threats: Grading",
      "quote": "reliability degrades most under formatting perturbations, to which graders are less robust than to semantic paraphrase",
      "evidence_links": {
        "33": [
          "R33-E02"
        ]
      }
    },
    {
      "id": "C48",
      "page": 15,
      "numbers": [
        26
      ],
      "marker": "[26]",
      "context_before": "ute procedure, so a grader could notice that the bundles fall into groups and could\nattempt to infer which group was treated. Nothing in the design bounds that effect. The\nannotation is arm-correlated: the three annotated runs had 12 of 17 candidates confirmed,\na lower share than the 7 of 8 of the one baseline run in that batch frozen normally, and\nwith different runs behind the two ratios neither figure isolates an effect of the annotation.\n     Subjects, seeds and one environment defect. Two subject systems and thirty\nseeded defects bound what generalizes, and seeded defects cannot be shown to represent\nthe faults real development produces ",
      "context_after": ". Both systems and the worker’s client are pub-\nlic open-source software, so their code and documentation may be in the worker model’s\ntraining data, ",
      "file": "07-threats.tex",
      "line": 52,
      "keys": [
        "bures2020injection"
      ],
      "section": "§7 Threats: Subjects",
      "quote": "seeded defects cannot be shown to represent the faults real development produces",
      "evidence_links": {
        "26": [
          "R26-E03"
        ]
      }
    },
    {
      "id": "C49",
      "page": 15,
      "numbers": [
        22
      ],
      "marker": "[22]",
      "context_before": "rrelated: the three annotated runs had 12 of 17 candidates confirmed,\na lower share than the 7 of 8 of the one baseline run in that batch frozen normally, and\nwith different runs behind the two ratios neither figure isolates an effect of the annotation.\n     Subjects, seeds and one environment defect. Two subject systems and thirty\nseeded defects bound what generalizes, and seeded defects cannot be shown to represent\nthe faults real development produces [26]. Both systems and the worker’s client are pub-\nlic open-source software, so their code and documentation may be in the worker model’s\ntraining data, a known confound in oracle evaluation ",
      "context_after": "; the design did not test whether\nprior knowledge of these systems interacts with the treatment. The PrestaShop storefront\nalso returned HTTP 404 for ",
      "file": "07-threats.tex",
      "line": 55,
      "keys": [
        "molinelli2025usefuloracles"
      ],
      "section": "§7 Threats: Subjects",
      "quote": "Both systems and the worker’s client are public open-source software, so their code and documentation may be in the worker model’s training data, a known confound in oracle evaluation",
      "evidence_links": {
        "22": [
          "R22-E01",
          "R22-E02"
        ]
      }
    }
  ]
}
