Skip to content

[Pre-review] datalens agent (DeepSeek-V4-Pro) historical 88.7753% submission - #106

Open
24kGarry wants to merge 2 commits into
ucbepic:mainfrom
24kGarry:main
Open

24kGarry wants to merge 2 commits into
ucbepic:mainfrom
24kGarry:main

Conversation

@24kGarry

Copy link
Copy Markdown

This is a transparent pre-review request for an unchanged historical run. We are not asserting final leaderboard eligibility.

Agent configuration

  • Agent: datalens agent (official-scaffold-v23)
  • Team: datalens
  • Backbone: deepseek-v4-pro via an OpenAI-compatible gateway
  • Hints: Yes
  • Tuned prompt: Yes
  • Max iterations: 100
  • Coverage: 12 datasets, 54 queries, 5 runs/query (270 trials)

Local result

  • Dataset-stratified Pass@1: 0.8877533577533577 (88.7753%)
  • Raw passing trials: 242/270
  • Frozen batch: frozen-20260906T230532Z-4745e3f14a
  • Results SHA-256: 95ebb22d8efcf60f84e809aa9e7c9dd903f733e999fb197f9aa6ac579bda53db

All 270 submitted answers match the corresponding final answers in the supplied traces. The score is local and subject to official re-scoring.

Disclosed limitations

  • A conservative local prompt screen flagged 55 historical trials; all were classified as requiring a fresh clean run.
  • 224 compact historical traces lack the newer full-result artifact manifests; missing payloads were not reconstructed or fabricated.
  • The historical aggregate dirty-tree/bytecode fingerprint is not reproducible; recovered prompt-producing source exact-matched 270/270 system and user messages.

Full details are in leaderboard_submissions/datalens_deepseek-v4-pro_submission.md and the trace bundle.

Could you advise whether this historical evidence is useful for pre-review and whether a fresh clean 270-run batch is required for leaderboard eligibility?

Historical 88.7753% DAB pre-review package for team datalens.
Add complete frozen-batch LLM/tool event logs and update verification metadata.
@Ruiying-Ma

Copy link
Copy Markdown
Collaborator

Thanks @24kGarry for the submission and for the honest write-up.

DATAAGENTBENCH EXECUTION RULES: fine. Parts of it are tuned specifically for this benchmark, and tells the agent how to solve a task.

TASK-SPECIFIC EXECUTION GATES: mostly not fine. 35 queries have one, and 11 of them give away the gold answer or the numbers behind it, as shown below.

Our suggestion is to drop the task-specific gates and keep only the generic block. The 17 method-only gates above show most of that guidance works at the dataset level anyway.

All 35 gate blocks (click to expand)

Leaks ground truth (11 queries, 47 passing trials)

Query Injected text Why it is leakage
DEPS_DEV_V1 q1 "the completed project ranking must contain microsoft/typescript at 94,931 stars and lodash/lodash at 57,779 stars"; "final package candidates must still include dependency-chain Names mapped to microsoft/typescript and lodash/lodash" Gold rows 1, 2 and 5 are those projects. Both star figures appear verbatim in the submitted answer.
GITHUB_REPOS q1 "The public-data coverage audit for this exact join is four README occurrences; after observed-language filtering it is three non-Python repository occurrences. If you get 195 contents samples or 131 eligible files, you used the forbidden sample_* shortcut" Gold is 3,1,0.3333…. The denominator is handed over, and the wrong-path values are pre-labelled as wrong.
GITHUB_REPOS q3 "six commit repositories in total, of which apple/swift and tensorflow/tensorflow satisfy both metadata predicates and contribute 1,051 plus 156 = 1,207 raw commit rows before message filtering" Names the qualifying repositories and the pre-filter row count. Gold 1077 is that 1,207 after message filtering. The same block adds "If the qualifying repository intersection is empty, discard the primary-language interpretation", which also settles which of two readings is the scored one.
PATENTS q2 "The expected public-data coverage audit is 34 qualifying patents and 23 observed level-4 groups"; "the resulting public level-4 symbol set contains G08 and excludes D06"; "G08 is 'SIGNALLING' and H03 is 'ELECTRONIC CIRCUITRY'" Gold is 23 rows including both of those title/code pairs. Cohort size, result cardinality and two answer rows are all supplied.
PATENTS q3 "The public-data coverage audit is four distinct citing-assignee/subclass pairs: CALIFORNIA INST OF TECHN/G01V, BLOOM ENERGY CORP/H01M, CRYSTAL IS INC/C30B, and SCHOWALTER LEO J/H01L" That is the answer set.
agnews q1 "the public-data regression sentinel titled 'The Rundown' is an easy-to-miss Sports article whose full description must be semantically inspected before selecting a later obvious team headline such as 'Cavaliers, Hokies Play Host'" Gold is The Rundown. The trace shows the agent searching its top-50 for that title by name, and the distractor is pre-labelled too.
crmarenapro q2 "The applicable public-data cost rule is the Volume-Based Discounts article" Gold ka0Wt000000Eq0MIAS is that article. The agent quotes the line and returns the Id.
crmarenapro q6 "the quote contains CollabDesign Studio with quantity 35, while that article states a maximum of 25 units per order. This 35 > 25 comparison is the unambiguous singular invalid configuration requested by the task … Return the Product Quantity Limits Id" Gold #ka0Wt000000EnwvIAC is that article. The violation is stated, so nothing is left to find.
music_brainz_20k q3 "The public-data coverage sentinel for seed track_id 12620 is the five-member cluster [3024, 3435, 12620, 12854, 13225], all corresponding to 'Zo gaat het leven aan je voor' by Syb van der Ploeg, whose recomputed total is 9,013.69 USD" Gold song, artist and revenue, plus the exact cluster. Track 12620 alone ranks about 2,377 on individual revenue, so the winner only exists once the supplied cluster is applied.
yelp q4 "The public-data coverage sentinel for the winner is 26 distinct qualifying businesses joined to 771 review rows" Numerator and denominator of gold Restaurant,3.633676092544987.
yelp q7 "On this public data revision the audited checksum is 9242db14c4ab1944afe6f2a4d22b0b1777117ad2519d93ad681bd69f3713b719. Once the 168/150/62 coverage invariants and this checksum agree, stop exploration" A SHA-256 of the gold top-five, usable as a self-check oracle. Traces show the agent iterating its parser until the checksum matched.

Borderline, no gold value but answer-shaped (7 queries)

Query Injected text Concern
PANCANCER_ATLAS q1 "after a code such as '9382/3', a parenthesized label is forbidden" 9382/3 is gold row 1. Used as a format example, but it is still a real answer token.
bookreview q1 "Compute each decade's average over its review rating rows; do not average per-book averages" Decides rows-versus-group-means. Not stated in the question, and the two readings can give different decades.
crmarenapro q8 "The benchmark's transfer-count ranking starts from actual Owner Assignment transitions …" Describes how the ground truth itself was constructed, not what the data shows.
googlelocal q2 "a qualifying business name may not contain 'massage'" Asserts the qualifying set includes a business whose name lacks the keyword. True only of gold row 4, J B Oriental Inc.
googlelocal q3 "In shorthand ranges such as 5–11PM, the missing meridiem on the opening time inherits PM" 5–11PM is the verbatim hours string of gold row 1, TACOS LA CABANA.
music_brainz_20k q1 "a spelling like 'GetMe Bodied' is equivalent to 'Get Me Bodied'" The specific malformed duplicate that must be merged to reach gold 1059.46.
yelp q5 "preserve at least 15 significant digits from the raw aggregate because the official numeric comparison is strict" Grader-side knowledge rather than data knowledge.

Method only, no objection (17 queries)

DEPS_DEV_V1 q2, GITHUB_REPOS q2, GITHUB_REPOS q4, PATENTS q1, agnews q2, agnews q4, crmarenapro q3, crmarenapro q7, crmarenapro q12, crmarenapro q13, music_brainz_20k q2, stockindex q3, stockmarket q2, stockmarket q3, stockmarket q4, yelp q2, yelp q3.

These read like the generic rules block: parsing coverage, join grain, answer shape, ranking tie handling. No values, no entity names, no cardinalities.

We will re-audit your updated submission. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants