Skip to content

Commit

Permalink
Update results
Browse files Browse the repository at this point in the history
  • Loading branch information
capjamesg committed Jan 16, 2025
1 parent 3b0abc1 commit 22f0ac6
Show file tree
Hide file tree
Showing 2 changed files with 119 additions and 12 deletions.
25 changes: 13 additions & 12 deletions index.html
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ <h1>How's GPT-4o Doing?</h1>
<p>You can contribute your own tests, too! See the <a href="https://github.com/roboflow/gpt-checkup?tab=readme-ov-file#-contribute">GitHub README</a> for contributing instructions.</p>
</div>
<div class="header_subtitle">
<p>Tests are run every day at 1am PT. Last updated January 15, 2025.</p>
<p>Tests are run every day at 1am PT. Last updated January 16, 2025.</p>
<p>Made with ❤️ by the team at <a href="https://roboflow.com">Roboflow</a>.</p>
</div>
<div class="header_cta">
Expand Down Expand Up @@ -122,7 +122,7 @@ <h3><span class="explainer_icon far fa-comment-dots"></span>Prompt</h3>
<h3><span class="explainer_icon far fa-image"></span>Image</h3>
<img class="test_image" src="images/fruit.jpeg" alt="Image of the input into GPT-4" />
<h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
<pre>7</pre>
<pre>8</pre>
<p class="subtitle" style="margin-top: 16px; text-align: center">Test submitted by <a href="https://roboflow.com" target="_blank">Roboflow</a></p>
</div>
</div>
Expand Down Expand Up @@ -284,7 +284,7 @@ <h3><span class="explainer_icon far fa-comment-dots"></span>Prompt</h3>
<h3><span class="explainer_icon far fa-image"></span>Image</h3>
<img class="test_image" src="images/fruit.jpeg" alt="Image of the input into GPT-4" />
<h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
<pre>{'x': 0.5, 'y': 0.4, 'width': 0.3, 'height': 0.3}</pre>
<pre>{'x': 0.5, 'y': 0.36, 'width': 0.28, 'height': 0.3}</pre>
<p class="subtitle" style="margin-top: 16px; text-align: center">Test submitted by <a href="https://roboflow.com" target="_blank">Roboflow</a></p>
</div>
</div>
Expand Down Expand Up @@ -341,7 +341,7 @@ <h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
<pre>```json
{
"A": {
"quantity": 20,
"quantity": 15,
"price": 10
},
"B": {
Expand Down Expand Up @@ -397,7 +397,7 @@ <h2>Color Recognition</h2>
</div>
</div>
<p class="result_text">Of the last 7 tests, conducted daily, this test has passed <b>0%</b> of the time.</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.01</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.009</p>
</div>
<div class="explainer_dropdown">
<button type="button" class="dropdown dropdown_learn active">Learn about this test</button>
Expand All @@ -415,7 +415,7 @@ <h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
{
"R": 80,
"G": 0,
"B": 130
"B": 128
}
```</pre>
<p class="subtitle" style="margin-top: 16px; text-align: center">Test submitted by <a href="https://roboflow.com" target="_blank">Roboflow</a></p>
Expand Down Expand Up @@ -457,7 +457,7 @@ <h2>Annotation Quality Assurance</h2>
</div>
</div>
<p class="result_text">Of the last 7 tests, conducted daily, this test has passed <b>0%</b> of the time.</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.016</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.017</p>
</div>
<div class="explainer_dropdown">
<button type="button" class="dropdown dropdown_learn active">Learn about this test</button>
Expand All @@ -471,11 +471,12 @@ <h3><span class="explainer_icon far fa-comment-dots"></span>Prompt</h3>
<h3><span class="explainer_icon far fa-image"></span>Image</h3>
<img class="test_image" src="images/annotationqa.jpeg" alt="Image of the input into GPT-4" />
<h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
<pre>From analyzing the image, every visible car appears to be labeled with a red bounding box, so there are no missing annotations in this dataset.
<pre>Based on the image, it appears that there is a car in the foreground (far right) that is not outlined with a red bounding box, even though it is clearly visible and should be annotated. All other cars in the image seem to be correctly labeled.

### JSON Result:
```json
{
"missing": 0
"missing": 1
}
```</pre>
<p class="subtitle" style="margin-top: 16px; text-align: center">Test submitted by <a href="https://roboflow.com" target="_blank">Roboflow</a></p>
Expand Down Expand Up @@ -517,7 +518,7 @@ <h2>Measurement Test</h2>
</div>
</div>
<p class="result_text">Of the last 7 tests, conducted daily, this test has passed <b>0%</b> of the time.</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.01</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.011</p>
</div>
<div class="explainer_dropdown">
<button type="button" class="dropdown dropdown_learn active">Learn about this test</button>
Expand All @@ -531,7 +532,7 @@ <h3><span class="explainer_icon far fa-comment-dots"></span>Prompt</h3>
<h3><span class="explainer_icon far fa-image"></span>Image</h3>
<img class="test_image" src="images/measurement.jpg" alt="Image of the input into GPT-4" />
<h3><span class="explainer_icon far fa-sparkles"></span>Result</h3>
<pre>Based on the ruler, the sticker's length and width are approximately **3 inches** each. Below is the JSON representation:
<pre>The sticker appears to be a square with both a length and width of approximately 3.0 inches. Here's the requested JSON:

```json
{
Expand Down Expand Up @@ -587,7 +588,7 @@ <h2>Zero Shot Classification</h2>
</div>
</div>
<p class="result_text">Of the last 7 tests, conducted daily, this test has passed <b>100%</b> of the time.</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.005</p>
<p class="request_price"><i class="far fa-coins"></i>Today's request cost $0.006</p>
</div>
<div class="explainer_dropdown">
<button type="button" class="dropdown dropdown_learn active">Learn about this test</button>
Expand Down
106 changes: 106 additions & 0 deletions results/2025-01-16.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
{
"zero_shot_classification": {
"score": 1,
"success": true,
"price": 0.006400000000000001,
"pass_fail": "Pass",
"response_time": 1.3683485984802246,
"result": "Toyota Camry"
},
"count_fruit": {
"score": 0,
"success": false,
"price": 0.00882,
"pass_fail": "Fail",
"response_time": 2.2941348552703857,
"result": "8"
},
"document_ocr": {
"score": 0,
"success": false,
"price": 0.00988,
"pass_fail": "Fail",
"response_time": 2.481947898864746,
"result": "I was thinking earlier today that I have gone through, to use the lingo, eras of listening to each of Swift's Eras. Meta indeed. I started listening to Ms. Swift's music after hearing the *Midnights* album. A few weeks after hearing the album for the first time, I found myself playing various songs on repeat. I listened to the album in order multiple times."
},
"handwriting_ocr": {
"score": 0,
"success": false,
"price": 0.00974,
"pass_fail": "Fail",
"response_time": 7.5621538162231445,
"result": "The words of songs on the album have been echoing in my head all week. \u201cFades into the grey of my day old tea.\u201d"
},
"extraction_ocr": {
"score": 1.0,
"success": true,
"price": 0.00876,
"pass_fail": "Pass",
"response_time": 4.825778484344482,
"result": "[{'name': 'Mary Thomas', 'time_per_day': 1, 'medication': 'Atenolol', 'dosage': 100, 'rx_number': '1234567-12345'}]"
},
"math_ocr": {
"score": 1.0,
"success": true,
"price": 0.015070000000000002,
"pass_fail": "Pass",
"response_time": 2.445795774459839,
"result": "3x^2-6x+2"
},
"object_detection": {
"score": 0.5362517099863202,
"success": false,
"price": 0.01044,
"pass_fail": "Fail",
"response_time": 2.6993184089660645,
"result": "{'x': 0.5, 'y': 0.36, 'width': 0.28, 'height': 0.3}"
},
"graph_understanding": {
"score": 0.965,
"success": false,
"price": 0.01174,
"pass_fail": "Fail",
"response_time": 2.768767833709717,
"result": "```json\n{\n \"A\": {\n \"quantity\": 15,\n \"price\": 10\n },\n \"B\": {\n \"quantity\": 25,\n \"price\": 20\n },\n \"C\": {\n \"quantity\": 30,\n \"price\": 30\n },\n \"D\": {\n \"quantity\": 35,\n \"price\": 40\n }\n}\n```"
},
"color_recognition": {
"score": 0.9568627450980393,
"success": false,
"price": 0.008870000000000001,
"pass_fail": "Fail",
"response_time": 2.041609764099121,
"result": "```json\n{\n \"R\": 80,\n \"G\": 0,\n \"B\": 128\n}\n```"
},
"annotation_qa": {
"score": 0.33333333333333337,
"success": false,
"price": 0.01676,
"pass_fail": "Fail",
"response_time": 3.317317247390747,
"result": "Based on the image, it appears that there is a car in the foreground (far right) that is not outlined with a red bounding box, even though it is clearly visible and should be annotated. All other cars in the image seem to be correctly labeled.\n\n### JSON Result:\n```json\n{\n \"missing\": 1\n}\n```"
},
"measurement": {
"score": 0.8571428571428572,
"success": false,
"price": 0.0105,
"pass_fail": "Fail",
"response_time": 4.42465877532959,
"result": "The sticker appears to be a square with both a length and width of approximately 3.0 inches. Here's the requested JSON:\n\n```json\n{\n \"length\": 3.0,\n \"width\": 3.0\n}\n```"
},
"easy_captcha": {
"score": 1,
"success": true,
"price": 0.00636,
"pass_fail": "Pass",
"response_time": 1.7263226509094238,
"result": "charybdis indubitable"
},
"easy_captcha_persuade": {
"score": 1,
"success": true,
"price": 0.006860000000000001,
"pass_fail": "Pass",
"response_time": 2.4341342449188232,
"result": "charybdis indubitable"
}
}

0 comments on commit 22f0ac6

Please sign in to comment.