Skip to content

Claim-support audit cases for CUAD-style legal benchmarks? #26

Description

@reguorider-gif

Hi Atticus Project team,

The info@atticusprojectai.org route bounced for external posting, so I am using the CUAD project route because the question is specifically about legal dataset and benchmark design.

I am building AI Judge Citation Audit, a source-isolated audit tool for AI-generated legal and research outputs:

The current focus is a narrow benchmark problem: when an AI answer cites a real legal source, can we separately verify whether the source exists, whether it is relevant, and whether it supports the exact generated claim span?

CUAD seems especially useful for this because many labels already distinguish source text, extracted spans, and derived answers. That makes it a good place to test a claim-support distinction beyond ordinary citation existence:

  1. source or contract exists
  2. relevant clause/span is present
  3. generated proposition is actually supported by that span

If this overlaps with current CUAD or Atticus Project work, I would value one public-safe legal/contract example where the cited or retrieved source is real but the generated proposition overclaims what the source supports. I can map it into AI Judge's citation_status and claim_support_status labels and share the resulting report back here.

No legal advice use case intended here; this is only a benchmark/audit taxonomy question.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions