Gemini 3.7 Flash vs 3.6: both wrote the same headline
Gemini 3.7 Flash (High) vs Gemini 3.6 Flash (High)

Based on 2 real-world tests
2 shared tests, none of them decided yet.
2 shared tests
One test has separated them on stated requirements. It is at the top of the chart below.
| Measure | ChatGPT | Gemini |
|---|---|---|
| Average score | — | — |
| Web app completion time | 7m 0s | 9m 0s |
| Community | 0% | 0% |
A blog page people finish: 46 KB vs 4.4 MBWebsites
ChatGPT vs Gemini: the rougher game is the better oneGames
Share of the requirements the brief actually stated, decided by reading the files each model produced. Sorted by the gap between the two, so the tests that separated them come first.
| Test | ChatGPT | Gemini |
|---|---|---|
| ChatGPT vs Gemini: the rougher game is the better one | 11/11 | 11/11 |
| A blog page people finish: 46 KB vs 4.4 MB | 7/7 | 6/7 |
A blog page people finish: 46 KB vs 4.4 MB
ChatGPT by 2m 0s7m 0s vs 9m 0s
Measured in each provider’s own web app, so these include network and interface behaviour. They are not model inference latency.
Gemini 3.7 Flash (High) vs Gemini 3.6 Flash (High)

ChatGPT vs Gemini

Gemini 3.1 Pro (High) vs Gemini 3.7 Flash (High)

ChatGPT vs Claude

Gemini vs DeepSeek

Gemini · 4 attempts
