# Study Finds Small Language Models Can Handle Most Home‑Automation Tasks

> Researchers evaluated nine models from 0.8 B to a frontier AI and showed that local models achieve similar accuracy to hosted ones for most functions.

Oossa · 2026-10-08 · https://oossa.com/en/study-finds-small-language-models-can-handle-most-home-automation-tasks

A new paper released on October 8, 2026 examines how a home‑automation system called Wactorz talks to language models. The system makes several kinds of calls – intent routing, action classification, device grounding, planning pipelines, and generating Python code. The authors tested nine models ranging from 0.8 billion parameters to a large hosted model. They used the system’s real production prompts on two Home Assistant installations, covering 280 real cases and 2 520 scored calls.

## What the numbers show

The results reveal that bigger models are not always better. For four of the five call sites, the best local model performed statistically the same as both the small hosted model and a frontier model. Only the code‑generation site showed a measurable gap, with the best local model scoring slightly worse (p = 0.039 vs. the small host and p = 0.002 vs. the frontier host). Overall accuracy with per‑call routing to the optimal local model was 91.8 %, compared with 95.4 % when every call used the most capable model, but without any extra compute cost. In a live user test, handling only the two generative sites on a host matched full hosting (39 correct out of 43) while using just 28 % of the spending.

## Safety quirks of tiny models

The benchmark also uncovered a safety issue at the actuation site. One 4 B model (Gemma4 E2B) mistakenly tried to control devices it didn’t own in 87.2 % of requests, while another model refused every request. These extreme behaviors illustrate that small models can resolve the accuracy‑refusal trade‑off in unpredictable ways.

## The facts

- The study evaluated nine language models from 0.8 B to a frontier hosted model.
- Testing used two real Home Assistant installations, covering 280 cases and 2 520 scored calls.
- Best local model achieved 91.8 % accuracy versus 95.4 % for the single hardest‑site model.
- Only code generation showed a significant difference (p = 0.039 vs. small host, p = 0.002 vs. frontier).
- Gemma4 E2B actuated on 87.2 % of requests for devices it did not own.

## Why it matters

For homeowners using open‑source automation, the findings mean they can run small, locally stored models for most tasks without paying for cloud hosting, keeping data private and costs low. However, they should be cautious with actuation commands, as some tiny models may behave unpredictably and either over‑act or refuse entirely.

## Sources & references

1. [Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System](https://arxiv.org/abs/2610.09021) – arXiv, 2026-10-08

Last updated: 2026-10-08
