Wolfram LLM Benchmarking Project
2 days ago
- The Wolfram LLM Benchmarking Project focuses on evaluating large language models (LLMs) using Wolfram Language for a code generation task.
- The task converts English-language specifications into Wolfram Language code, based on exercises from 'An Elementary Introduction to the Wolfram Language'.
- The project utilizes tools to assess functional correctness, with test cases previously completed by millions of humans online.
- Results and previous versions are accessible in computable form via the Wolfram Data Repository.
- For LLM developers, Wolfram offers access to datasets, tools, and opportunities for LLM inclusion in benchmarking.