← Back to Blog

Embedded Cat Sound Detection: Offline Meow Recognition SDK with Cry DiscriminationNEW

The Value of Cat Sound Detection: Scenarios and Workflow

Cats are among the few pets that "talk" to their humans: short meows when hungry, long yowls under stress or in heat cycles, early-morning wake-up calls. For pet products, cat sound detection upgrades the device from "seeing your pet" to "understanding your pet" — with the microphones already in the BOM.

Scenario families: pet-monitoring cameras, smart feeders and litter boxes, boarding facilities with multi-unit monitoring, and — a fast-growing combination — pet-and-baby households, where cat detection and cry detection work as a pair.

Indoor scene and device forms
Cameras, feeders and multi-area monitoring in one engine

The Signature — and the Most Underrated Challenge

Cat signature vs baby cry and birdsong
Pitch glides vs cry harmonics vs high-frequency chirps

Below is a real cat sound recording — waveform (top) and Mel spectrogram (bottom). The short bursts and mid-frequency pitch-glide structure are clearly visible:

Real cat sound sample: waveform and spectrum
Real cat sound sample: waveform (top) + Mel spectrogram (bottom)

Audio samples (real recordings — press play):

🔊 Cat sound (real sample)

🔊 Baby cry — compare its similarity to the meow

Meows — a distinctive "rise-then-fall" pitch glide, 0.3–1 second per call

Yowls — longer and louder, arriving in stretches during stress or heat

• Clearly distinct from birdsong (high-frequency chirps) and dog barks (lower mid-frequency bursts in series)

The underrated challenge: meows sit remarkably close to baby cries in the spectrum — mid-frequency energy, clear pitch glides, rich harmonics. Generic audio models routinely confuse the two: baby monitors report meows as crying, pet cameras mistake cries for meows. Cat detection therefore needs dedicated bidirectional discrimination — which is exactly what makes the "cat + cry" combined offering work in pet-and-baby homes.

How Recognition Works: Preprocessing → Features → Algorithms

1. Events as units — output onset/offset and intensity for each vocal segment, not per-frame labels

2. Pitch-glide verification — the rising-falling pitch contour of a meow is unique evidence separating it from steady sounds

3. Dedicated discrimination against cries and speech — trained with targeted negative samples in both directions to cut cross false alarms

On-device architecture
Capture, processing, engine, output, application

Alone-Time Care Workflow

Cat sound monitoring workflow
Monitoring, event detection, activity statistics, threshold push and behavior report

While the owner is away or at night: event detection → activity statistics (count / duration / time-of-day) → push an alert with live video when the concern threshold is exceeded, otherwise log into the behavior report. Over time, the data reveals the classic dawn/dusk activity rhythm.

From Vocal Data to Health Insight

A single meow is an event; months of data are a signal. Long-term vocalization patterns become a personal health record — not a medical device, but a behavior reference that spots trends owners would miss.

Behavior baseline
Four-week baseline with deviation detection (illustrative)

Behavior baseline: during the first 1–2 weeks the system learns this cat's personal baseline — average daily meow count, active hours, typical call duration. After that, deviations become meaningful: sustained doubling above baseline, unusual night calls, or abnormally long calls. The principle: daily fluctuation is normal; sustained deviation is the signal.

Multi-metric correlation
Meowing, eating and litter-box trends over four weeks (illustrative)

Multi-metric correlation: one metric can be noisy; the report therefore correlates three streams — meowing, eating, litter-box activity. When several drift in the same window, it deserves attention.

Life-stage differences
Kitten, adult and senior vocal patterns (illustrative)

Life stages: kittens call frequently in short bursts; adult patterns are regular and easy to baseline; senior cats may show more night-time long calls — watch the trend, and consult a vet when pronounced.

Health monitoring workflow
From daily capture to graded reminders and long-term records

Graded response: daily capture → baseline learning → deviation detection → graded reminders (gentle note / suggest observation / suggest consulting a vet) → owner action → long-term record. Positioning statement: a behavior-analysis and care-reference tool — it does not constitute a medical diagnosis; when sustained anomalies appear, combine the data with daily observation and consult a professional veterinarian.

Weekly report interface
The weekly report and four reading cues

Device Forms and Integration

Four device forms
Cameras, feeders, litter boxes, boarding facilities

Cameras prioritize low-latency push; feeders prioritize low power plus motor-noise robustness; litter boxes emphasize stability under self-noise; boarding facilities need multi-channel concurrency — all sharing one SDK.

Selection framework
Four-dimension evaluation framework

Evaluate along four dimensions: recognition capability (accuracy, false alarms, cry discrimination), resource footprint (0.2–1 MB INT8, CPU under 50 MHz), platform fit (ARM/MIPS/x86_64, NPU pre-adaptation), and engineering support (samples, integration days, licensing, OTA).

Integration flow
Six steps from pilot scenario to volume licensing

Typical rhythm: pilot scenario → validate with your own audio online → integrate the C API (days) → joint debugging on real hardware → two-week field trial → volume licensing with OTA updates.

Multi-device coordination
Devices detect locally; events aggregate to the user app

Multiple devices detect locally; events and statistics are shared and surfaced in one daily report — all inference on-device.

Accuracy and Performance

Item
Spec
Accuracy
95.5%
False alarm rate
<2%
Model size
0.2–1 MB (INT8)
Inference latency
Configurable to platform resources
Sample rate
16 kHz
Platforms
ARM Linux / MIPS / x86_64; SVP / Magik pre-adapted

Note: performance figures are based on internal test environments; actual results depend on hardware and deployment scenarios.

Platform and Hardware Requirements

From 100 MHz-class chips upward; the model shares the SoC with the device's main workload; battery-powered feeders and litter boxes are supported by the low-power event architecture.

C API and Embedded Integration

The cat sound detection library exposes a concise streaming C API: the caller just keeps feeding 16 kHz mono PCM; framing, Mel preprocessing and model inference run internally, and frame-level probabilities are aggregated by the alarm strategy into event callbacks.

Full interface declaration (cat_detect.h):

cat_detect.hc
/**
 * cat_detect.h — 猫叫声识别统一接口
 *
 * 封装 Mel 预处理 + 推理引擎 + 报警策略, 内部模型消费线程处理音频。
 * 与录音模块 (audio_capture.h) 相互独立: 调用者自行决定音频来源
 * (录音回调 / wav 文件 / 网络流), 通过 cat_detect_feed 送入, 数据任意大小。
 *
 * 用法 (实时录音模式):
 *   cat_detect_t *d = cat_detect_create(mgk_path, NULL, NULL);
 *   cat_detect_set_listener(d, on_frame, on_onset, on_offset, NULL);
 *   cat_detect_start(d);                          // 启动内部模型消费线程
 *   audio_capture_start(rec, capture_cb, d);        // 录音回调里调 cat_detect_feed
 *   ...
 *   cat_detect_stop(d);                           // 排空缓冲, 停止线程
 *   cat_detect_destroy(d);
 *
 * 用法 (wav 文件模式):
 *   cat_detect_t *d = cat_detect_create(mgk_path, NULL, NULL);
 *   cat_detect_set_listener(d, on_frame, on_onset, on_offset, NULL);
 *   cat_detect_start(d);
 *   循环读文件: cat_detect_feed(d, pcm, n);       // 任意数据大小
 *   cat_detect_stop(d);
 *   cat_detect_destroy(d);
 */

#ifndef CAT_DETECT_H
#define CAT_DETECT_H

#include <stdint.h>

#ifdef __cplusplus
extern "C" {
#endif

/* 识别事件 (报警策略输出, 用于事件结束回调) */
typedef struct {
    float start_time;       /* 事件开始时间 (秒) */
    float end_time;         /* 事件结束时间 (秒) */
    float confidence;       /* 事件置信度 */
    float max_confidence;   /* 事件内最大帧置信度 */
    int   frame_count;      /* 事件持续帧数 */
} cat_detect_event_t;

/* 帧级回调: 每帧识别结果 (模型线程内执行) */
typedef void (*cat_detect_frame_cb_t)(float cat_prob, float timestamp,
                                        void *user_data);

/* 事件开始回调: 策略判定猫叫事件开始, 只有开始时间 */
typedef void (*cat_detect_onset_cb_t)(float start_time, void *user_data);

/* 事件结束回调: 策略判定猫叫事件结束 (或停止识别时未结束的事件), 完整事件信息 */
typedef void (*cat_detect_offset_cb_t)(const cat_detect_event_t *event,
                                         void *user_data);

typedef struct cat_detect_s cat_detect_t;

/* 创建/销毁; alarm_name/alarm_params 可传 NULL (用默认策略及参数) */
cat_detect_t *cat_detect_create(const char *mgk_path,      /* 模型文件路径 (必填) */
                                    const char *alarm_name,    /* 报警策略名, NULL=默认 */
                                    const char *alarm_params); /* 策略参数 key=val,key=val, NULL=默认 */
void cat_detect_destroy(cat_detect_t *det);

/* 设置事件回调 (create 后调用, 也可在运行中调整); 不需要的回调传 NULL */
void cat_detect_set_listener(cat_detect_t *det,
                               cat_detect_frame_cb_t  on_frame,
                               cat_detect_onset_cb_t  on_onset,
                               cat_detect_offset_cb_t on_offset,
                               void *user_data);

/* 启动/停止识别: 启动内部模型消费线程 / 排空缓冲后停止线程 */
int  cat_detect_start(cat_detect_t *det);
void cat_detect_stop(cat_detect_t *det);
int  cat_detect_is_running(cat_detect_t *det);

/* 设置事件识别策略 (可在运行中调整) */
int cat_detect_set_alarm(cat_detect_t *det, const char *alarm_name,
                           const char *alarm_params);

/* 送入 PCM 数据 (16bit 单声道 16kHz), 线程安全, 任意数据大小 */
int cat_detect_feed(cat_detect_t *det, const int16_t *pcm, int num_samples);

#ifdef __cplusplus
}
#endif

#endif /* CAT_DETECT_H */

A minimal WAV-inference demo (excerpt; the full file ships at src/catDetect/c/cat_demo.c):

cat_demo.cc
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include "cat_detect.h"

#define DEFAULT_MGK  "cat_detect_v7.mgk"  /* 模型文件 */
#define DEFAULT_WAV  "test_cat.wav"       /* 16kHz 单声道 16bit PCM */
#define FEED_CHUNK   16000                  /* 每次送入 1 秒音频 */

/* 帧级回调: 每帧输出猫叫声概率 (约 1 秒一帧) */
static void on_frame(float cat_prob, float timestamp, void *user_data)
{
    (void)user_data;
    printf("%8.2fs  cat=%.4f\n", timestamp, cat_prob);
}

/* 事件开始回调: 策略判定猫叫事件开始 */
static void on_onset(float start_time, void *user_data)
{
    (void)user_data;
    printf("[EVENT] cat start at %.2fs\n", start_time);
}

/* 事件结束回调: 应用层可据此做后续统计或分级响应 */
static void on_offset(const cat_detect_event_t *ev, void *user_data)
{
    (void)user_data;
    printf("[EVENT] cat end at %.2fs (dur=%.2fs, conf=%.3f, frames=%d)\n",
           ev->end_time, ev->end_time - ev->start_time,
           ev->confidence, ev->frame_count);
}

static int read_wav_pcm(const char *path, int16_t **pcm, int *n, int *sr);  /* 完整实现见源文件 */

int main(void)
{
    cat_detect_t *det;
    int16_t *pcm = NULL;
    int num_samples = 0, sample_rate = 0;
    int pos;

    if (read_wav_pcm(DEFAULT_WAV, &pcm, &num_samples, &sample_rate) != 0)
        return 1;

    /* 1. 创建识别器: 模型文件 + 默认报警策略 (NULL) */
    det = cat_detect_create(DEFAULT_MGK, NULL, NULL);
    if (!det) return 1;

    /* 2. 注册回调 (均为可选) */
    cat_detect_set_listener(det, on_frame, on_onset, on_offset, NULL);

    /* 3. 启动内部模型消费线程 */
    cat_detect_start(det);

    /* 4. 分块送入 PCM; 实时录音时改在录音回调里 feed */
    for (pos = 0; pos < num_samples; pos += FEED_CHUNK) {
        int n = num_samples - pos;
        if (n > FEED_CHUNK) n = FEED_CHUNK;
        cat_detect_feed(det, pcm + pos, n);
    }

    /* 5. 停止并销毁 */
    cat_detect_stop(det);
    cat_detect_destroy(det);
    free(pcm);
    return 0;
}

Build and run:

buildbash
$(CC) cat_demo.c -I. -L. -lcatdetect -lpthread -lm -o cat_demo
./cat_demo

Conclusion

Cat sound detection turns a pet device from a gadget into a care companion — understanding what the cat is expressing right now, and spotting long-term trends. An online trial with full technical support is available.