Spaces:
Running
Benchmark Feature Request: Models with Similar Writing Styles
There is always some inherent bias in evaluation metrics, and one important source of that bias is writing style. Different models naturally produce different styles (purple vs beige), and everyone have preferences for one over another. It would be useful to expose writing style as its own set of metrics and make it possible to compare models with similar writing styles, rather than relying on aggregate scores alone.
In addition, each model profile could include a Writing Style section that summarizes the model's overall prose and storytelling tendencies in plain language (e.g., concise vs. verbose, dialogue-heavy vs. descriptive, and what I'm personally looking for: hentai terminology richness), along with recommendations for models that exhibit similar writing styles.