Research manuscript

Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes

Hao Li1, Jinye Zhang1, Bobo Li2, Mong-Li Lee2, Wynne Hsu2, Zheng Wang1, Hao Fei3,*, Min Zhang4

1Wuhan University 2National University of Singapore 3University of Oxford 4Harbin Institute of Technology (Shenzhen) *Corresponding author

Motivation

Existing affective benchmarks often evaluate local targets such as affective states or causes. TRACE-Bench asks five complementary questions about one episode, from the conditions that give rise to an emotion, through the regulation that reshapes its expression, to the consequences it produces for the subject or other participants.

Affective Blueprint

Condition, Affect, and Effect describe a subject-centered affective episode, with each stage grounded in a corresponding theory.

Select a task to highlight its supplied and queried constructs.

QueriedSupplied* When annotated

③ Effect ℰ

Mental Effect Affective Effect Physical Effect

② Affect 𝒜

Manifestation Aman

Facial Bodily Verbal Vocal
f(·)

State Asta

Object Polarity Category (coarse & fine) Intensity

Regulation Areg

Target State Source State Tactic Goal

First-order

Second-order

① Condition 𝒞

Physical Condition

Environment Characters Event

Mental Condition

Belief Desire Intention

Cause Ccau

External Subject Internal

TRACE-Bench

Each task takes a multimodal social video with aligned subtitles as input. The tasks differ in which Blueprint constructs are supplied as known and which are queried, covering within-stage reasoning (T1–T2), cross-stage reasoning (T3–T4), and full-chain reconstruction (T5).

Structured QA pairs
3,746
Videos
646
Tasks
5
Video scenes 646 videos
Area chart of paper-reported scene estimates: Home 240, Outdoor 128, Office 90, Institutional 47, School 39, Restaurant 38, Medical 34, Leisure 30.
Social relationships 646 videos
Paper-reported relationship estimates: Romantic 161, Authority 141, Professional 109, Family 103, Friends 90, Strangers 42.
QA pairs by task 3,746 total
Proportional task segments. QA counts: T1 919, T2 401, T3 795, T4 1123, T5 508.
Emotion categories T1 annotations
12 emotion categories: Anger 294, Surprise 205, Sadness 203, Fear 146, Happiness 138, Disgust 118, Guilt 60, Frustration 49, Anxiety 48, Embarrassment 44, Relief 40, Neutral 11.
Regulation tactics 401 T2 pairs
Ring chart of tactics: Suppression 189, Substitution 98, None 76, Amplification 29, Fabrication 9.
Effect types 1,123 T4 pairs
Three-column comparison: Physical 452, Affective 385, Mental 286.

Statistics reported in the paper; scene and relationship counts are estimates. Emotion categories count labels, not QA pairs. Areas and segments show proportions; bars compare counts from zero.

Polarity, intensity, and cause structure
Polarity 919 T1 pairs
Negative 734, Positive 142, Mixed 23, Neutral 20.
Intensity 919 T1 pairs
Medium 422, High 415, Low 76, None 6.
Cause structure 795 T3 pairs
External only 409; internal and external 386.
T1 Affect recognition
Identify what the person feels and how they express it in the scene.
T2 Regulation decoding
Given the affect category, intensity, and person it is directed toward, infer how the emotion is regulated and cite supporting cues.
T3 Cause reasoning
Given the person's affect, identify the triggering event and, where annotated, the beliefs, desires, or intentions behind it.
T4 Effect reasoning
Given the person's affect, infer its mental, emotional, or behavioral consequence for the specified recipient.
T5 Full-chain reconstruction
Reconstruct the causes, affect, and requested consequences of the episode.
Benchmark Construction
  1. 1

    Source collection

    Social clips from Movie-clips, EQ4You, and Social-IQ are screened by an MLLM and verified during annotation.

  2. 2

    Chain annotation

    An MLLM drafts subject-centered chains; annotators reject, correct, add, and verify them against the video.

  3. 3

    QA annotation

    Templates convert each verified chain into task-specific QA pairs, followed by final human validation.

Evaluation Analysis

Open-source, closed-source, and affect-specialized MLLMs were evaluated on TRACE-Bench, alongside two diagnostic studies of their outputs.

01

Human vs. Model Performance

GPT-5 (thinking) scores 20.40 points below humans.

On the same 50 questions per task, GPT-5 (thinking) averages 58.75, compared with 79.15 for humans. TRACER scores 63.58 on this subset.

02

Model Family Comparison

The best emotion-focused model averages 21.47, compared with 57.96 for GPT-5 (thinking).

Each dot is one baseline's average over the five tasks on the full benchmark.

03

State Attribution

Models sometimes mistake what people show for what they actually feel.

In the tested moments where feelings and expressions differ, GPT-5 makes this mistake in 19% of descriptions and Qwen3-VL in 27%. Giving the models an analysis of each character's situation improves correct identification of feelings from 50% to 60% and from 40% to 51%, respectively.

04

Unsupported-Content Audit

For claims about what people do next, 17.3% are made up by GPT-5, compared with 3.3% by TRACER.

Full-chain outputs are checked against video frames and subtitles in External Reality, Manifestation, and Physical Effect. TRACER reduces the aggregate fabrication rate from 11.0% to 4.7%.

TRACER

Each derivation step couples a natural-language inference with explicit premise citations drawn from factual observations, cognitive appraisals, and established upstream conclusions. The linked steps form a traceable graph of intermediate and target conclusions.

Explore the premise links in Fig. 5(b)

Derivation Example

Select a conclusion to see the premises it cites.

Showing all premise links.

Adapted from Fig. 5(b) of the paper.

Main Results

Differences are significant on every task (p < 0.05, bootstrap test, Holm-corrected). The largest gains appear in regulation decoding and cause reasoning.

Scores on the main evaluation set. Avg. averages the five tasks.

All models and the human-matched scores
Main results on TRACE-Bench (Table 2 of the paper)
ModelT1 Rec.T2 Reg.T3 CauseT4 EffectT5 FullAvg.
Affect-specialized
Emotion-LLaMA0.455.22*2.26†0.31†0.20†1.69
AffectGPT34.574.64†17.99*15.79*20.87*18.77
Emotion-Qwen29.45*24.24*21.04*13.4919.13*21.47
Open-source general-purpose
LLaVA-OneVision-7B35.1323.5125.284.115.47†18.70
LLaVA-NeXT-Video-32B25.8927.0127.13*23.6113.80*23.49
MiniCPM-V-4.537.3029.1943.0428.2729.4333.45
GLM-4.6V-Flash-9B40.8028.8542.3028.8132.3634.62
LLaVA-OneVision-70B42.7426.7347.3033.0334.6236.88
Qwen3-Omni-30B46.2833.1751.6939.8141.7842.55
InternVL3.5-38B45.3528.0546.1632.1937.0737.76
Qwen3-VL-32B50.9735.4960.4244.3850.1648.28
Closed-source general-purpose
GPT-5 (non-thinking)56.1337.2670.7452.3658.2254.94
Gemini-3-Pro54.3240.5570.7451.4856.0554.63
GPT-5 (thinking)56.6141.5677.1153.7560.7957.96
TRACER (ours)61.6749.5384.5559.4763.5963.76

* marks cells with substantial parsing failures, and † cells where most outputs could not be parsed.

Human and model performance on the same 50 questions per task (Appendix A.7)
Model / HumanT1T2T3T4T5Avg.
Human83.6072.3285.5578.9075.4079.15
GPT-5 (thinking)57.9741.8775.8057.6160.4858.75
GPT-5 (non-thinking)55.2136.0068.1261.0659.0955.89
Gemini-3-Pro55.5941.5368.0060.1954.1855.90
Qwen3-VL-32B50.5634.2663.6649.0050.6849.63
Qwen3-Omni-30B49.3229.3754.3242.6841.5543.45
InternVL3.5-38B48.0028.6444.0534.9435.1638.16
TRACER (ours)60.7849.9184.3960.2162.5963.58

Human scores average two participants.

Experimental Analysis

Evidence Ablation

Joint subtitles and frames outperform either channel alone on all five tasks for both models. Doubling GPT-5's frames yields only small, inconsistent gains.

Fig. 7. Q: question and supplied fields. Frames: GPT-5 8, Qwen 60; Both (16f): GPT-5 with subtitles and 16 frames.

Model Family Preferences

Affect-specialized models favor T1–T2; closed-source models favor T3 and T5. Open-source general-purpose models are more balanced but weaker on T5.

Fig. 6(b). Family- and task-centered preferences, not absolute rankings.

T1 · Recognition Errors

Models confuse similar emotions, tend to predict medium intensity, and often miss that positive and negative feelings can coexist.

Fig. 8. True → predicted labels; widths show within-class proportions.

T2 · Regulation Decoding

TRACER is better than GPT-5 (thinking) at recognizing when people suppress or amplify their emotions.

Fig. 9. Reference rows, predicted columns, row percentages. Fabrication omitted.

T3 · Cause Reasoning

Internal Driver scores lower across all evaluated models. Thinking narrows GPT-5's gap; TRACER leads on both components but retains the same imbalance.

Fig. 10. Dashed line: equal scores. Arrow: GPT-5 with thinking.

T4 · Effect Channels

Physical Effect is hardest for general-purpose baselines. TRACER leads on Mental and Physical Effect, while Affective Effect remains close to GPT-5 (thinking).

Fig. 11. Other-directed instances only; dashed lines separate model families and TRACER.

Case Studies

Each panel shows the task input, selected observations and appraisal readings, the construct outputs, and a comparison with the reference and two baselines.

T1 Affect recognition
T2 Regulation decoding
T3 Cause reasoning
T4 Effect reasoning

Fig. 12 from the paper. Click a panel to enlarge. Text is condensed from the source records; T1 comes from a separate illustrative run.

BibTeX

@misc{li2026trace,
  title  = {Cognition-Oriented Emotion Tracing from Causes
            to Consequences in Real-World Social Scenes},
  author = {Li, Hao and Zhang, Jinye and Li, Bobo and
            Lee, Mong-Li and Hsu, Wynne and Wang, Zheng and
            Fei, Hao and Zhang, Min},
  year   = {2026},
  eprint = {2610.11410},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url    = {https://arxiv.org/abs/2610.11410}
}