In short
Can a general-purpose AI model, with no training, tell from a low-resolution thermal sensor (heat only, no ordinary camera) whether someone fell? We asked Google's Gemini exactly as our research prototype does, on 438 hand-located staged falls from a public thermal dataset and on every clip the prototype produced over 76 hours of ordinary life in a lived-in room, and we compared it with detectors trained the conventional way.
- No false alarms in ordinary life. Over 76 hours of everyday activity in a lived-in room the AI made no false fall calls, and it called none of the 499 laboratory clips of people walking, sitting, lying on the floor and getting up a fall.
- It catches most falls it can see, not all. It caught 69% of the staged falls of people it had never seen in a public thermal dataset, and 4 of 9 staged falls the camera could see in the room.
- One sentence of instructions made a large difference. A cautionary line telling the AI that a person who goes out of view has not fallen cost about 14 points of fall sensitivity on the same laboratory falls (69% with it, 83% without it).
- How many pixels it needs. Fewer pixels, fewer falls caught: 69% at 32 by 24 pixels, 48% at 16 by 12, and 34% at 8 by 6, the size of the cheapest sensors.
- Every answer can be checked. When the AI calls a fall it says when it happened, within a second of the real moment in 99% of cases, so a person can verify it in seconds.
Questions this paper answers
Can AI detect a fall from a thermal camera?
Partly. With no training, Google's Gemini 3.8 Flash caught 69% of the staged falls of people it had never seen in a public 32 by 24 pixel thermal dataset, and 4 of 9 staged falls a 160 by 120 thermal camera could see in a lived-in room. Detectors trained on the same laboratory data caught nearly every laboratory fall but raised 14 and 24 false alarms a day at home.
How many false alarms does AI fall detection on thermal video give?
In 76 hours of ordinary life in a lived-in room it made no false fall calls (95% upper bound 1.2 a day), and it called none of the 499 laboratory clips of people walking, sitting, lying on the floor and getting up a fall.
Why does an AI miss falls on thermal video?
Mostly because of one sentence in its instructions. Told that a person who goes out of view has not fallen, it declined falls that ended at the edge of the frame or behind furniture, often after describing the fall itself. Without that sentence it caught 83% of the laboratory falls instead of 69%, and 9 of the 9 visible home falls instead of 4, but it also called deliberate lie-downs on the floor a fall.
How many pixels does a thermal sensor need for AI fall detection?
Every halving of resolution cost falls: 69% of the laboratory falls were caught at 32 by 24 pixels, 48% at 16 by 12 and 34% at 8 by 6, the size of the cheapest thermal arrays, with few added false alarms.
Does a thermal sensor protect privacy compared with a camera?
At the resolutions in this study (32 by 24 and 160 by 120 pixels) thermal video shows a warm body and its posture, not a face. This paper does not measure privacy; our next study compares thermal and normal video of the same moments, including how much each reveals about who the person is.
Does the AI give the same answer twice?
Not always. Sending the identical request again at temperature 0 changed 4% of the verdicts, mostly on borderline falls. When it does call a fall it says when it happened, so a person can check the clip in seconds.
Abstract
Thermal cameras can watch over older people living alone without capturing faces, and multimodal large language models (LLMs) can judge a short video without task-specific training. We evaluate a production multimodal LLM, Gemini 3.8 Flash at temperature 0, as the decision stage of an event-triggered fall screener, used exactly as a deployed prototype uses it: one request per short thermal clip, with its sound track when there is one, and its predecessor Gemini 3.7 Flash on the same requests. On 438 staged falls and 563 non-fall windows that we located by hand in a public 32x24 thermal dataset, the deployed request called 69% of the held-out subjects' falls (67% over all subjects; bed falls excluded) and none of 499 non-fall windows, including people lying on the floor and getting up. Over 76 h of ordinary life in an occupied room, recorded at 160x120 with sound, it called none of 395 clips a fall (95% upper bound 1.2 false alarms per day) and caught 4 of the 9 staged falls the camera could see. One cautionary sentence in the prompt ("a person who goes out of view ... is not a fall") accounted for many of the misses: removing it raised sensitivity on the held-out subjects from 69% to 83% on 3.8 Flash and from 80% to 92% on Gemini 3.7 Flash, and in the home from 4 to 9 of the 9 visible falls, with no false call in ordinary living but at the price of calling deliberate lie-downs a fall. Sensitivity collapsed at 8x6 pixels (34%); identical requests disagreed on 4% of windows; and when the model called a fall, its explanation placed it within a second of the annotated landing 99% of the time. Two conventional detectors trained on the laboratory data caught nearly every held-out laboratory fall but raised 14.2 and 23.9 false alarms per day in the home. We release the frame-level fall annotations of the public dataset and every model answer on it.
How to cite
Junca, M., & Junca, M. (2026). Asking a Multimodal LLM Whether Someone Fell: Zero-Shot Screening of Low-Resolution Thermal Video in a Laboratory and an Occupied Home. Senecta Inc. https://senecta.ai/research/thermal-fall-detection-llm/
@techreport{junca2026thermal,
title = {Asking a Multimodal {LLM} Whether Someone Fell: Zero-Shot Screening of Low-Resolution Thermal Video in a Laboratory and an Occupied Home},
author = {Junca, Marcos and Junca, Mauro},
institution = {Senecta Inc.},
year = {2026},
month = oct,
url = {https://senecta.ai/research/thermal-fall-detection-llm/}
}
Honest scope
- The falls were staged, by younger adults, in one laboratory and one room. Real falls of older people are different.
- The room recordings come from a research prototype, not the product's final sensor, mounting or software.
- These are research results, not a promise of how often a product will catch a fall. No system catches every fall.
Senecta makes the pods described in this research, so this is our own study, not an independent evaluation. The paper describes the methods, the data and every limitation we know of.