Search
SuperBench
a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on their performance in different aspects such as natural language understanding, reasoning, and generalization.
...moreSuperLim
a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such as argumentation analysis, semantic similarity, and textual entailment.
...moreTAT-DQA
a large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly in financial reports.
...moreTAT-QA
a large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.
...moreVisualWebArena
a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.
We-Math
a benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
WHOOPS!
a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.
...moreDeepSeek-Math-7B
LLM application: DeepSeek-Math-7B
DeepSeek-Coder-1.3|6.7|7|33B
LLM application: DeepSeek-Coder-1.3|6.7|7|33B
DeepSeek-VL-1.3|7B
LLM application: DeepSeek-VL-1.3|7B
DeepSeek-MoE-16B
LLM application: DeepSeek-MoE-16B
DeepSeek-Coder-v2-16|236B-MOE
LLM application: DeepSeek-Coder-v2-16|236B-MOE
DeepSeek-V2.5
LLM application: DeepSeek-V2.5
Qwen-1.8B|7B|14B|72B
LLM application: Qwen-1.8B|7B|14B|72B
Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B
LLM application: Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B
Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B
LLM application: Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B
Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B
LLM application: Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B
CodeQwen1.5-7B
LLM application: CodeQwen1.5-7B
Qwen2.5-Coder-1.5B|7B|32B
LLM application: Qwen2.5-Coder-1.5B|7B|32B
Qwen2-Math-1.5B|7B|72B
LLM application: Qwen2-Math-1.5B|7B|72B