NoticeGuard benchmark

Rubric-consistency on synthetic cases

NoticeGuard versus a plain LLM given the same documents, two prompt styles, three runs each.

Rubric-consistency on synthetic cases. Not legal validation. Expected verdicts were labelled by hand against a published rubric. The baseline is the same model NoticeGuard uses for extraction. Baseline outputs are shown unedited — click any cell to view them.

Loading results…

Raw output