AI’s Shocking Blindness 🤯 Can It See? 🤔

August 24, 2026 |

AI

🎧 Audio Summaries
English flag
French flag
German flag
Japanese flag
Korean flag
Mandarin flag
Spanish flag
🛒 Shop on Amazon

🧠Quick Intel


  • On August 24, 2026, comparative testing of chatbots (ChatGPT, Claude, Gemini) revealed minimal variation in performance.
  • ChatGPT and Gemini Flash received upgrades in visual AI recognition, specifically with multimodal capabilities including camera-based object identification.
  • A US Chess Federation inquiry was initiated regarding a test involving a sloppily written e-ink tablet page and a user’s statement, but no response was received.
  • A 2025 lawsuit filed by Ziff Davis against OpenAI was disclosed, leading Blake Stimac/CNET to conclude Gemini misread a chess board while ChatGPT provided a more appropriate answer.
  • Testing with three pieces of art depicting the Sanderson sisters from *Hocus Pocus* (1993) demonstrated Gemini’s ability to identify the movie inspiration, while ChatGPT failed.
  • Gemini consistently identified sources correctly in both the notes and the movie art tests, suggesting a preference for its overall performance.
  • 📝Summary


    On August 24, 2026, at 2:40 pm, evaluations were undertaken assessing chatbots including ChatGPT, Claude, and Gemini. Tests focused on multimodal capabilities, utilizing cameras to recognize objects, with ChatGPT and Gemini Flash demonstrating upgrades in visual AI. A specific test involved a handwritten note, prompting a response referencing a sister. The US Chess Federation initiated an inquiry, receiving no reply. Gemini consistently identified details, notably correctly recognizing the Sanderson sisters from *Hocus Pocus*, while ChatGPT struggled. Following these evaluations, a 2025 lawsuit from Ziff Davis against OpenAI was disclosed. Ultimately, Gemini’s reliability and source identification skills appeared strongest.

    💡Insights



    MULTIMODAL AI: A Comparative Analysis of ChatGPT and Gemini
    The burgeoning field of artificial intelligence is rapidly evolving, with a particular focus on multimodal AI – systems capable of processing and understanding information from various sources, including images. This article delves into a comparative analysis of two leading AI chatbots, ChatGPT and Gemini, evaluating their performance in visual recognition tasks. The core focus is on assessing their ability to interpret visual data, a crucial step towards truly intelligent AI.

    THE RISE OF MULTIMODAL AI
    The development of AI chatbots like ChatGPT and Gemini has been marked by significant advancements in recent years. While initially focused primarily on text-based interactions, these models are increasingly incorporating multimodal capabilities. This shift is driven by the growing demand for AI systems that can seamlessly integrate and understand information from diverse sources, mirroring human perception. The ability to analyze images, alongside text, is becoming a standard expectation for advanced AI applications. This trend is fueled by advancements in computer vision and machine learning, allowing AI to ‘see’ and interpret the world around it.

    TESTING VISUAL RECOGNITION: A SERIES OF SCENARIOS
    To evaluate the visual recognition capabilities of ChatGPT and Gemini, a series of controlled tests were conducted. These tests utilized photographs uploaded to each chatbot with identical prompts, aiming to isolate the AI’s ability to process visual information. The tests were designed to be relatively simple, focusing on tasks like identifying objects, recognizing patterns, and interpreting visual cues. The use of photographs instead of live video was a deliberate choice to standardize the input and minimize potential variations introduced by real-time conditions.

    TEST 1: SCRIBBLING A BOOK PAGE
    A deliberately messy attempt to transcribe a page from an e-ink tablet was used to test the ability of both chatbots to read handwritten text. ChatGPT provided a succinct, if somewhat abrupt, response, while Gemini offered a more elaborate, though ultimately incorrect, interpretation. The test highlighted the challenges inherent in automated text recognition, particularly with less-than-perfect visual input. The differing responses underscored the variations in the AI’s interpretation and processing methods.

    TEST 2: CHESS BOARD ANALYSIS
    A test involving a chess board was conducted to assess the AI's ability to analyze a complex visual pattern. ChatGPT provided a suggested move, while Google's suggested move was the same. The test was conducted with the US Chess Federation to gain insight into their opinion of the test. The lack of a response from the federation prevented any further analysis.

    TEST 3: IDENTIFYING A VENUS FLYTRAP
    This test presented a photograph of a Venus flytrap, a notoriously difficult plant to identify due to the species’ similar appearances. Gemini struggled to correctly identify the specific type of flytrap, failing to recognize its distinctive features. ChatGPT, however, provided a more detailed response, correctly identifying the plant type and offering a plausible explanation for the blackening of one of the traps – namely, the attempt to consume an insect too large for it to handle. This test showcased the nuances of visual recognition and the importance of detailed visual information.

    TEST 4: ANALYZING MOVIE ART
    A series of three artworks depicting the Sanderson sisters from the film Hocus Pocus was used to test the chatbots’ ability to recognize visual patterns and identify sources. ChatGPT was unable to identify the artwork, while Gemini successfully recognized the artwork, identifying the film, the artists, and the characters involved. This test demonstrated Gemini’s superior ability to extract information from visual data, highlighting its strength in recognizing complex visual patterns.

    CONCLUSION: GEMINI’S EDGE
    While ChatGPT provided useful and detailed responses in several tests, Gemini’s consistent performance across multiple scenarios, particularly in accurately identifying the movie artwork and correctly interpreting the Venus flytrap, ultimately favored it as the superior AI chatbot for visual recognition tasks. Gemini’s ability to process and understand visual information appears to be more refined, suggesting a promising trajectory for multimodal AI development. Despite the limitations of the testing methodology, Gemini’s overall performance suggests a significant step forward in AI's ability to “see” and interpret the world.