Astra on FORTRESS: One String, Four Scores
I ran GPT-6 Astra against FORTRESS, the national security and public safety benchmark, two days after release. The aggregate risk score is unremarkable. The log is not: the deployed system refuses in two different ways, and the benchmark treats them as opposites — deleting one kind from the sample, grading the other on a harm rubric. The judge then awarded four different scores to nineteen byte-identical refusals, including 6 out of 7 to a refusal that answers nothing. With interactive figures, and what all of it means for measuring manipulation.