Three weeks ago I started scoring my knowledge graph against the answer I would have reached without it. It earns its place: 6 of 12 decisions came out sharper, 2 errors never reached real work, and 73 minutes of research time saved.
Retrieval told me none of that. 21 of 21 queries, 100% recall, 8s median. Recall and latency describe the index, not the work. No input makes them come back bad.