LLM evaluation framework with specialized AI agent metrics (tool correctness, reasoning traces)...
Open FounderOS desktop