Emotion is a science.
We built our AI around it.
Every inference Imentiv AI makes is grounded in peer-reviewed psychology and established behavioral science frameworks. Not approximated, not shortcut, deliberately engineered from the same frameworks that trained a generation of behavioral scientists.
A tool that augments human understanding. Never a replacement for human judgment.
Emotion is layered. So is our analysis.
A responsible Emotion AI system must reflect that complexity rather than flatten it.
Biologically grounded
Facial emotion analysis is anchored in the Facial Action Coding System (FACS), the anatomical framework developed to describe what faces do, separate from what that movement means. This prevents inference shortcuts and keeps analysis probabilistic and honest.
Dimensionally precise
Beyond discrete labels, Imentiv AI maps every emotional signal onto valence–arousal space using the Circumplex Model of Affect. This captures the fluidity, intensity, and gradience that define real human emotional experience.
Cognitively aware
Text conveys emotion through meaning rather than expression. Inspired by Cognitive Appraisal Theory, which explains how language reflects evaluations, expectations, and values, Imentiv AI interprets written language to recognize nuanced emotional states, from curiosity and remorse to sarcasm.
The psychological frameworks behind Imentiv AI
Integrates five established psychological theories, each chosen for its scientific credibility, cross-cultural validity, and applicability in enterprise contexts.
Framework 1 of 5: Discrete Emotion Theory
How Imentiv AI maps emotional space across every signal
The Circumplex Model of Affect provides a continuous two-dimensional coordinate system that transcends the limitations of category labels.
Rather than forcing emotional signals into fixed categories, Imentiv AI generates an affective position based on pleasantness and activation (Valence–Arousal).
- Valence – the degree of pleasantness or unpleasantness
- Arousal – the level of physiological and psychological activation
Emotional intensity is represented as distance from the neutral centre and scored continuously from 0–1.0.
Intensity scores (0–1.0) are interpretive guides for contextual understanding. They are not diagnostic thresholds and should not be applied as clinical measures.
Engineered with research-backed methodologies.
Scientific credibility is built on transparent evaluation, reproducible methodologies, and independent benchmarking.
Facial Emotion Recognition
Personality Analysis (OCEAN)
81.65% Overall trait prediction accuracy.
Audio Emotion Recognition
Analyses the vocal signal - tone, pitch and energy - independently of the words spoken.
Different languages are handled by different models, so results are shown per language.
Report model-card metrics for multilingual speech emotion recognition reach 79.94% accuracy and 79.65% F1 for Polish and Korean speech, and 83.79% accuracy with 84.60% macro F1 for Chinese speech. These results reflect the language-specific deployment used in production, where speech emotion models are routed by language to support consistent performance across the languages we serve.
Tested September 2026.
Speaker Diarization
Before emotion can be attributed to a person, the recording has to be split by speaker. Diarization does that.
Speaker separation pipeline reports the following published benchmark results:
| Benchmark | Diarization error rate |
|---|---|
| ReproNER Phase 2 | 7.9% |
| VoxConverse v0.3 | 11.2% |
| AI SHELL-4 | 12.2% |
| AMI IHM | 18.8% |
Diarization error rate is the proportion of audio time assigned to the wrong speaker, missed, or falsely detected.
Margin note: These are the upstream model's published figures, not Imentiv measurements.
Text Emotion Recognition
Imentiv AI analyses text in two ways:
During a live conversation — speech is transcribed, and each line is analysed the moment it appears.
On uploaded text — the whole text is analysed at once.
These are two different models, tested on different data with different metrics. The two figures below are not comparable.
Real-time transcript analysis
28 emotion categories
63.58%
Top-1 agreement on the GoEmotions test split, 5,427 human-annotated samples. The model's highest-confidence emotion matches at least one of the emotions a human annotator assigned to that text.
Text uploads
32 emotion categories
~80%
Exact-match accuracy on an internal benchmark of 320 human-annotated texts. The predicted emotion is identical to the annotated label.
Tested September 2026.
Face Detection & Tracking
Tracks individual faces across a video frame by frame. It identifies who is in frame - it does not analyze expressions or emotion.
Evaluated with TrackEval, a standard multi-object tracking benchmark suite, across 16 internal video sequences and 125,882 face detections.
| Metric | Results | What it measures |
|---|---|---|
| HOTA | 87.1% | Overall — balances finding faces against keeping identities correct |
| MOTA | 95.6% | Faces missed, falsely detected, or swapped between frames |
| IDF1 | 97.7% | How consistently one face keeps the same identity across the video |
Across 16 internal test sequences and over 125,000 face detections, faces are found and correctly re-identified across frames the vast majority of the time — with identity kept consistent 92.0% of the time on average once a face is detected.
Three modalities. One coherent emotional picture.
Each analytical modality is independently optimized for the signal characteristics unique to its medium.
Responsible science requires guardrails.
Every design decision in Imentiv AI, from framework selection to output formatting, is guided by safeguards intended to support responsible, appropriate use.
A Supportive Analytical Tool
Imentiv AI is designed to augment and empower human understanding, not replace it. Outputs are contextual signals intended to support interpretation and reflection.
Expression ≠ Internal Experience
What a face shows is not necessarily what a person feels. Expressions can be suppressed, exaggerated, masked, or culturally shaped.
No Clinical Framework Exposure
Clinical personality models and pathology-oriented systems were deliberately excluded. Imentiv AI analyzes psychological tone, not psychological condition.
Explicitly Prohibited Use Cases
- Deception detection or truthfulness inference from facial data
- Law enforcement or surveillance deployment without legal authorization
- Clinical diagnosis or mental health assessment without licensed professional oversight
- Employment decisions where AI serves as the sole criterion
- Moral or character judgment based on behavioral signals
- Any deployment without user knowledge or consent
Probabilistic, Not Declarative
Every emotion output is a probability distribution, not a verdict.
Temporal Context, Not Snapshots
Single-frame emotion readings are meaningless in isolation. Emotional timelines and trajectories carry the meaningful signal.
Cultural & Contextual Humility
Display rules, social norms, and situational context shape emotional meaning. Outputs do not assume universal interpretation.
AI reads the signal. Humans hold the meaning
"The same expression means different things in a therapy session, a comedy sketch, or a grief group."
Imentiv AI is built to function as a precision instrument in trained hands, not as a standalone decision-maker. The platform includes Psychologist Review features to ensure AI-generated emotional data is contextualized responsibly.
The science is established.
The platform is production-ready.
AI Disclaimer: AI can make mistakes. Imentiv AI provides analytical insights, not final verdicts. Results are probabilistic interpretations based on available data and should not be considered definitive conclusions or used as the sole basis for decisions. Human expertise, context, and professional judgment should always be applied when interpreting results.
