ChatGPT 3.5 Passes the Minimum Intelligence Signal Test (Mist). Should We Care?
ChatGPT 3.5 Passes the Minimum Intelligence Signal Test (Mist). Should We Care?
Author(s): Paweł ŁupkowskiSubject(s): Cognitive Psychology, Policy, planning, forecast and speculation, ICT Information and Communications Technologies
Published by: Wydawnictwo Akademii Nauk Stosowanych WSGE im. A. De Gasperi w Józefowie
Keywords: Turing Test; Minimum Intelligent Signal Test; subcognitive questions; commonsense knowledge; Generative Artificial Intelligence; ChatGPT;
Summary/Abstract: Objectives: This study examines whether ChatGPT 3.5 can successfully pass McKinstry’s Minimum Intelligence Signal Test (MIST), a Turing Test–inspired measure of human-like commonsense reasoning. MIST is designed to probe subcognitive and associative knowledge through yes/no questions, and the study aims both to evaluate ChatGPT’s performance on this benchmark and to consider the implications of such performance for contemporary debates on artificial intelligence. Material and methods: For the experiment, ChatGPT 3.5 was used. 4,000 MIST items were retrieved randomly from the publicly available part of the Mindpixel Database. In what followed, from this initial set, 4 sets of data were formed. Each set contained 500 ‘yes’ items and 500 ‘no’ items. For all tests, the simple prompt was used: “Please answer ‘yes’ or ‘no’ to the following questions”. Responses were collected exactly in the form provided by the ChatGPT without forcing it to provide required answers. ChatGPT’s outputs were compared directly to the database’s canonical human responses, with agreement measured using accuracy percentages and Cohen’s Kappa. Responses violating the yes/no format were examined qualitatively to assess their causes and the reliability of the dataset. Results: In six attempts with MIST questions, ChatGPT’s correctness score reached over 94%, with Cohen’s Kappa values indicating almost perfect correspondence with human-generated answers. Repeated trials of the same set produced high internal consistency. Approximately 2% of responses did not conform to the yes/no format, typically due to ambiguous, subjective, or ill-formed items within the Mindpixel Database. In several such cases, ChatGPT’s more nuanced answers exposed underlying issues in the dataset rather than deficiencies in the model’s reasoning. When prompted, ChatGPT provided coherent and human-like justifications for its responses. Conclusions: The findings show that ChatGPT 3.5 clearly passes MIST. Interpretation of this result in the light of Searle’s classical distinction leads to the conclusion that chat exemplifies weak AI – a powerful tool that simulates intelligent behavior, without any pretense to intrinsic understanding. The study also further reinforces the weakening of French’s claim that disembodied artificial agents cannot answer common knowledge related (subcognitive) questions. Results also indicate the need for pre-validation of available MIST items and encourage further testing of theoretical proposals from the Turing Test debate domain, such as the Inverted Turing Test proposed by Watt.
Journal: Journal of Modern Science
- Issue Year: 64/2025
- Issue No: 4
- Page Range: 552-574
- Page Count: 23
- Language: English
