Data Science Wire

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

arXiv cs.AI22h4 min read

arXiv:2609.13725v1 Announce Type: new Abstract: An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normal

Read the full story at arXiv cs.AI

More in MLOps / LLMOps