File size: 1,794 Bytes
21d9130
 
 
 
 
 
 
 
 
 
3d08f1a
 
216a944
3d08f1a
216a944
4569cb2
216a944
4569cb2
216a944
4569cb2
216a944
4569cb2
cf71611
216a944
cf71611
7e65b3f
4569cb2
966f39f
3d08f1a
 
4569cb2
7f391d2
 
4569cb2
03508cf
4569cb2
7e65b3f
4569cb2
3d08f1a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
---
language:
- en
tags:
- testing
- llm
- rp
- discussion
---

# Why? What? TL;DR?

Various tests on, well, various LLM.

# Available Tests

## LLM Drawing Test (2026 ongoing)

This test is meant to evaluate models in a difficult task requiring a competency in spatial awareness, image recognition, tools calls, and creativity. In a new empty chat, the model is asked to copy the provided image to the best of its abilities. Model is then evaluated on accuracy of function calls (fail rate) and the resemblance to the original drawing.

- [Results](https://huggingface.co/SerialKicked/ModelTestingBed/discussions/4)


## DoggoEval (Done)

The goal of this test, featuring a dog (Rex) and his owner (EsKa), is to determine if a model is good at obeying a system prompt and character card. The trick being that dogs can't talk, but LLM love to.

- [Results and discussions are hosted in this thread](https://huggingface.co/SerialKicked/ModelTestingBed/discussions/1) ([old thread here](https://huggingface.co/LWDCLS/LLM-Discussions/discussions/13))
- [Files, cards and settings can be found here](https://huggingface.co/SerialKicked/ModelTestingBed/tree/main/DoggoEval)
- TODO: Charts and screenshots




# Limitations 

I'm testing for things I'm interested in. I do not pretend any of this is very scientific or accurate: as much as I try to reduce the amount of variables, a small LLM is still a small LLM at the end of the day. The results for other seeds, or with the smallest of change, are bound to give very different results. 

I usually give the different models I'm testing a fair shake in a more casual settings. I regen tons of outputs with random seeds, and while there are (large) variations, it tends to even out to the results shown in testing. Otherwise I'll make a note of it.