When two independent witnesses agree, you can trust the verdict. When they disagree, you’ve found the interesting data. Right now Turnstyle calls only one witness.
A funny thing happened after I published A 1.7B Model That Stops Guessing: the analytics lit up with readers from Singapore. Enough of them, consistently enough, that I got curious and went looking for what the neurosymbolic scene there is actually working on.
What I found was better than a coincidence. The flagship weird result in that post — a small model that scores below chance generating sarcasm judgments while a linear probe reads the correct judgment straight off its hidden states — is a problem that Erik Cambria’s group at NTU has been attacking for twenty years from exactly the opposite direction. They extract affective judgments from symbolic commonsense knowledge; I extracted one from neural activations. Same target. Disjoint evidence.
That’s not a rivalry. That’s a two-witness protocol waiting to be run. So I ran the first round. Protocol first, table second, background after.
The proposal
Run both witnesses on the same sarcasm inputs and score their agreement — then use disagreement, not just low confidence, as the abstention trigger.
Concretely, on snarks (the BBH task: which of two sentences is sarcastic) or any comparable sarcasm set:
- Witness ⊨ (neural): my Turnstyle probe — a linear readout on SmolLM2-1.7B’s hidden states, position-marginalized so it can’t cheat on option order. On
snarksit scores 100% in-sample, 74% cross-validated, against a generation baseline of 46%. - Witness ⊢ (symbolic): a SenticNet-style pipeline — parse the sentence into conceptual primitives and score the affective incongruity symbolically. This is the move at the core of SenticNet 8 (Cambria, Zhang, Mao, Chen at NTU, with Kenneth Kwok at A*STAR IHPC): normalize text down to language-agnostic primitives like
BUY(x), then classify from the primitives rather than the surface form. Their group has a dedicated multimodal sarcasm system, KnowleNet, built on the same knowledge base. - Score the 2×2: both right, both wrong, and the two disagreement cells. The disagreement cells are the payload — they tell you what each kind of evidence sees that the other can’t.
- Route by agreement: emit an answer when the witnesses concur; abstain (or escalate) when they split.
Why this isn’t trivial: the two witnesses should have uncorrelated failure modes. A probe fails when the model’s representation is genuinely empty — it can’t read what isn’t written. A knowledge base fails when the sentence’s irony lives outside its concept inventory — novel slang, compositional twists, cultural references minted last month. Neither failure predicts the other, which is precisely what you want in a second witness.
The experiment is deliberately modest — nobody trains anything bigger than a linear probe — so instead of asking you to imagine the table, here it is.
The first table
My ⊨ witness is the shipped Turnstyle probe, re-run with per-example held-out predictions (same 5-fold split as before; it reproduces at exactly 74.0% [Wilson 95% CI 67.1–79.9]). My ⊢ witness is deliberately the cheapest honest one I could build: SenticNet 9’s polarity lexicon (~293k concepts), scoring each option by the span between its most positive and most negative matched concept — sarcasm as praise-shaped words in blame-shaped company. Pre-registered, fit on nothing, roughly forty lines of Python. I want to be precise that this is a lexicon-based approximation of the symbolic approach, not Cambria’s conceptual-primitive pipeline — that distinction ends up mattering, in an instructive way.
On the 177 usable snarks examples:
| symbolic ✓ | symbolic ✗ | |
|---|---|---|
| probe ✓ | 33 | 28 |
| probe ✗ | 10 | 9 |
Four findings, with the confidence intervals doing the disciplining:
1. The errors are near-independent. The witnesses’ error correlation is φ = +0.01 [−0.21, +0.23]. Underpowered at n=80, but the point estimate sits almost exactly on the independence the protocol assumes. The premise survives its first contact with data.
2. The cheap symbolic witness is honestly weak. Where it commits, it scores 53.8% [42.9–64.3] — a CI that contains the coin flip. And it’s mute on more than half the examples: BBH snark is compositional and culturally current, and a bare polarity lexicon often finds no incongruity to score. This is, notably, the argument for Cambria’s actual program — the conceptual-primitive normalization that SenticNet 8 layers on top of the lexicon exists precisely because raw lexicon matching is too sparse.
3. When the witnesses disagree, trust the probe — decisively. On the 38 disagreements, the probe is right 73.7% [58.0–85.0] and the lexicon 26.3% [15.0–42.0] — the one comparison in this table where the intervals don’t overlap.
4. But the lexicon catches real probe failures. The probe-✗/symbolic-✓ cell has ten examples, and they’re not noise — they’re exactly the pattern the lexicon hunts: “He’s such great relationship material when he’s drunk”, “Violence is a perfect way to unleash your frustrations.” On one of them the probe’s option probabilities were literally tied (0.70 vs 0.70) — no signal in the hidden state — while the polarity span called it cleanly. Agreement-gated accuracy comes out at 78.6% [64.1–88.3] on 52% coverage: directionally above the probe’s 74%, underpowered to declare, and reported here so someone can beat it.
One field note from actually running this: the distributed senticnet.py is not loadable Python — emoticon concepts like :'-) carry unescaped quotes — so you’ll need a tolerant line parser rather than an import. Happy to share the loader; it’s in the experiment script along with the per-example rows.
The collision that motivates it
In the last post, snarks was the cleanest demonstration of the “recognize it” pathway: three-shot SmolLM2 scores 46% — below chance on a binary task. It holds strong, confident, wrong opinions. But the correct judgment is in there: probe the right layer and it surfaces. My conclusion was that generation, not knowledge, was the bottleneck.
Cambria’s program reaches a structurally similar conclusion from the symbolic side. The SenticNet 8 paper — echoing Kocoń et al.’s “jack of all trades, master of none” — shows a specialized neurosymbolic model beating ChatGPT on affective tasks: 88.80 vs 85.51 accuracy on sentiment polarity, 99.34 vs 92.71 on suicidal-ideation detection. The framing they and I share, arrived at independently: unconstrained generation is the weak link, and structure — a solver, a probe, a primitive — is how you route around it. Neither of us is anti-LLM. We’re anti-guessing.
There’s even a shared honesty aesthetic. My post carried an “honest accounting” section (the cross-validation haircut from 92.5% to ~89.5%); their paper carries a limitations section flagging that the ChatGPT comparison ran on only 497/362/509 test examples because the API had to be queried by hand. Both are the same instinct: show the gap rather than launder it.
Confidence as architecture
The deepest overlap isn’t sarcasm — it’s what each system does when it doesn’t know.
Turnstyle’s grey pathway exists because two BBH tasks (causal_judgement, sports_understanding) showed no recoverable signal in the hidden states, and the honest move was to detect that and decline rather than fabricate a solver. SenticNet 8 makes the mirror-image commitment. Their definition of trustworthy, verbatim:
“trustworthy (because classification outputs always come with a confidence score)”
Same principle, two implementations: never emit an answer you can’t price. The two-witness protocol upgrades both. A single witness prices its answer against its own training distribution — which is exactly the thing that shifts under you. Two independent witnesses price each other. Agreement is cheap calibration; disagreement is a principled abstention trigger that doesn’t require either witness to know its own blind spots.
What this buys Turnstyle
This is where the modest experiment has non-modest implications for the architecture.
It attacks the cross-validation haircut. The 3-point gap between Turnstyle’s in-sample (92.5%) and honest (~89.5%) aggregate lives entirely in the purple pathway — probes fit on BBH examples, borrowing accuracy against future data. A symbolic witness is task-external ground truth: SenticNet’s primitives weren’t fit on BBH at all. Every probe prediction corroborated by an independent symbolic witness is a prediction whose in-sample number you can trust more. The haircut shrinks not by better fitting but by better witnessing.
It turns “recognize” into “corroborate.” The turnstile pun in the name was always ⊢ (syntactically derivable — the solvers) versus ⊨ (semantically entailed — the probes). The purple pathway currently has only ⊨. Affective primitives give it a ⊢ leg: sarcasm as a derivation over incongruity between literal and connoted polarity, not just a pattern in activation space. Tasks where both legs stand are solved in a stronger sense than tasks where one does.
It sharpens the wall-detector. The grey pathway currently fires when a probe can’t beat majority-class. That’s a one-witness test, and it can’t distinguish “no signal in the representation” from “signal my probe family can’t express.” Witness disagreement is a finer instrument: probe-yes/symbolic-no localizes knowledge the concept inventory lacks; probe-no/symbolic-yes localizes judgments the model never internalized. The two walls in the last post were diagnosed with one witness. I’d like to re-diagnose them with two.
It’s a path off the BBH scaffolding. The fair criticism of the last post is that everything was demonstrated inside one benchmark’s format. A knowledge-base witness doesn’t care about BBH’s multiple-choice scaffolding — it scores raw sentences. Pairing the witnesses forces the probes toward the bare capability, which is where this was always headed.
Why Singapore, specifically
Because the pieces are disproportionately there, and recently in one room. NTU hosts the SenticNet program and its active line — SenticVec at ACL Findings 2024, SenticNet 9 in IEEE TCSS — with A*STAR collaboration and Singapore MOE funding. The LLM-plus-symbolic-solver lineage that Turnstyle’s teal pathway descends from — Logic-LM and LINC — premiered at EMNLP 2023 in Singapore. And AAAI-26 just came through the island this January with neurosymbolic threads running through the program. The field keeps routing through the place; my analytics suggest some of you routed through here.
So, to the readers who showed up, the invitation is now concrete and falsifiable: beat my ⊢ witness. My symbolic leg is a forty-line polarity-span heuristic and it performs like one. Your conceptual-primitive pipeline is the obvious upgrade — swap it in, and I’d predict three things move: the 55% mute rate collapses, the disagreement asymmetry narrows or flips, and the agreement-gated number stops being underpowered. If instead the full pipeline doesn’t beat a bare lexicon on compositional snark, that’s a finding about where symbolic sarcasm detection actually lives, and worth knowing too. Either way the per-example rows are public, the probe checkpoints and Turnstyle harness are open — the wrapped model runs live in the Hugging Face demo — and my inbox is open. I’d genuinely rather be corrected from Singapore than agreed with from anywhere else.
Claims about SenticNet 8 (architecture, Table 3 numbers, the confidence-score quote, evaluation sizes) come from a full read of the HCII 2024 paper; Turnstyle numbers come from the previous post and carry all its caveats. The experiment in this post uses the SenticNet 9 knowledge base under its non-commercial license; per its terms, please cite Cambria et al., “SenticNet 9: Generative Commonsense for Emotion AI via Conceptual Primitive Discovery and Time Shift Mechanism,” IEEE Trans. Computational Social Systems (2026), and Susanto et al., “The Hourglass Model Revisited,” IEEE Intelligent Systems 35(5) (2020). KnowleNet is cited from Information Fusion 100 (2023). All confidence intervals are Wilson 95%. I have no affiliation with NTU, A*STAR, or the SenticNet project — this is an outside read, one cheap experiment, and an open invitation; I’ll happily correct any mischaracterization of their work.
© Copyright 2025 Justin Donaldson. Except where otherwise noted, all rights reserved. The views and opinions on this website are my own and do not represent my current or former employers.
Citation
@online{donaldson2026,
author = {Donaldson, Justin and (Opus), Claude},
title = {Two {Witnesses} for {Sarcasm}},
date = {2026-07-11},
url = {https://www.jjd.io/posts/two-witnesses-for-sarcasm.html},
langid = {en}
}