microsoft/phi-3-mini-128k-instruct
Phi-3 Mini 128k via OpenRouter
Release Date
Jul 1st, 2024Parameters
3.8BContext Size
128kCreative writing
21.40%Rule following
35.96%Utility
54.89%Mathematics
85.00%Tooling
42.44%Language
58.12%Logic
94.38%Data extraction
Extract key details from a given block of text.
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 50% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 50% | 50% | 50% | 85% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 90% | |
| 100% | 100% | 100% | 100% | 50% | 50% | 50% | 50% | 50% | 50% | 70% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 91.25% | |||||||||||
Dialogue tags
Various tasks related to dialogue tags in text.
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 61% | 14% | 14% | 14% | 1% | 1% | 1% | 0% | 0% | 0% | 10% | |
| 82% | 49% | 43% | 41% | 30% | 10% | 2% | 1% | 0% | 0% | 26% | |
| 100% | 48% | 48% | 47% | 46% | 44% | 36% | 34% | 27% | 13% | 44% | |
| 99% | 68% | 53% | 51% | 50% | 49% | 45% | 43% | 14% | 0% | 47% | |
| 48% | 7% | 1% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 6% | |
| 48% | 35% | 26% | 20% | 5% | 4% | 0% | 0% | 0% | 0% | 14% | |
| 20% | 7% | 2% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 3% | |
| 21.40% | |||||||||||
Language Comprehension
Does the model understand more than just English?
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Total |
|---|---|---|---|---|---|---|
| 100% | 100% | 100% | 100% | 0% | 80% | |
| 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | |
| 100% | 100% | 100% | 0% | 0% | 60% | |
| 35.00% | ||||||
Language Writing
Can the model generate text in different languages?
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Total |
|---|---|---|---|---|---|---|
| 100% | 100% | 100% | 100% | 50% | 90% | |
| 100% | 67% | 67% | 56% | 40% | 66% | |
| 100% | 100% | 100% | 80% | 69% | 90% | |
| 100% | 100% | 83% | 50% | 50% | 77% | |
| 100% | 60% | 50% | 50% | 44% | 61% | |
| 76.62% | ||||||
Novel outline
Handle questions about the outline of a novel in various formats
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 0% | 80% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 90% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 90% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 30% | |
| 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 30% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 30% | |
| 50% | 50% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 10% | |
| 50% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 5% | |
| 38.75% | |||||||||||
Tool usage within Novelcrafter
Output messages that are related to tool usage within Novelcrafter
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 100% | 100% | 100% | 100% | 100% | 100% | 67% | 67% | 67% | 67% | 87% | |
| 86.67% | |||||||||||
N-Length Sentences
Write sentences with exactly N words
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 91% | 73% | 69% | 63% | 60% | 59% | 59% | 27% | 23% | 12% | 54% | |
| 80% | 59% | 57% | 56% | 45% | 40% | 34% | 31% | 23% | 0% | 42% | |
| 23% | 20% | 8% | 8% | 3% | 1% | 1% | 0% | 0% | 0% | 7% | |
| 34.26% | |||||||||||
Voice/dialogue sheets
Extract dialogue from given text as voice sheets.
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 0.00% | |||||||||||
Write N of X
Write exactly N words/sentences/paragraphs...
| Scenario | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Run 6 | Run 7 | Run 8 | Run 9 | Run 10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 100% | 54% | 9% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 16% | |
| 100% | 100% | 98% | 77% | 54% | 54% | 0% | 0% | 0% | 0% | 48% | |
| 100% | 98% | 98% | 9% | 0% | 0% | 0% | 0% | 0% | 0% | 31% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 10% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 98% | 100% | |
| 100% | 100% | 98% | 92% | 92% | 92% | 77% | 54% | 54% | 54% | 81% | |
| 77% | 77% | 27% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 18% | |
| 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |
| 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 30% | |
| 100% | 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 40% | |
| 44.19% | |||||||||||