IROS 2026 · Pittsburgh

Touch2Know Dexterous In-Hand Liquid Sensing via Mutual-Capacitance Fingerprints

What's inside the bottle? Just grasp it.
Ruixiang Deng1 Zhegong Shangguan2† Yang Hu1 Haozheng Bai1 Tingcheng Li3 Angelo Cangelosi2† IEEE Senior Member Wuqiang Yang1,* IEEE Fellow
1Manchester Tashan Joint Laboratory (UK) of Touch Sensors for Domestic Robots, Department of Electrical and Electronic Engineering, The University of Manchester, U.K.
2Cognitive Robotics Lab, Department of Computer Science, The University of Manchester, U.K.
3School of Electronic and Information Engineering, Suzhou University of Science and Technology, China
*Corresponding author: wuqiang.yang@manchester.ac.uk  ·  †Supported by the ERC eTALK Project (EP/Y029534/1)
IROS 2026 Pittsburgh · COROLAB · The University of Manchester · Tashan · IROS contributed paper — Manchester Tashan Joint Laboratory (UK) of Touch Sensors for Domestic Robots × Cognitive Robotics Lab
95.07%
Liquid acc.
Balanced split
90.00%
Liquid acc.
Day-4 holdout
99.34%
Container acc.
Balanced split
98.16%
Container acc.
Day-4 holdout
19 liquid classes · 4 container types · 4 days · 3,040 grasp trials
Abstract

Robots and people often need to know what liquid is inside a sealed bottle before acting, but vision can fail for opaque or visually similar contents, and continuous cameras may be undesirable in shared spaces. We present Touch2Know, a dexterous-hand sensing system that identifies liquid type and container type through the container wall during normal in-hand grasping, without opening the bottle. Touch2Know integrates five low-profile flex-PCB electrodes with guarded and shielded routing on a multi-finger hand, and measures multi-channel mutual capacitance to reduce common-mode drift and grounding sensitivity. An event-driven adaptive grasp controller stabilizes contact and limits over-compression across PET and glass bottles. For learning, we propose CapTF (Capacitive Temporal Fusion), a CNN–Transformer model that jointly predicts (i) liquid class and (ii) container type from 4-channel capacitive time series. We collect a dataset spanning 19 liquid classes (11 household liquids plus salt-water and sugar-water concentration series) and 4 container types over four days. The task is fine-grained, including visually ambiguous cases and liquids with closely related dielectric/conductive responses through container walls. With per-trial baseline-subtracted signals, CapTF achieves 95.07% / 99.34% liquid/container accuracy under a balanced split and 90.00% / 98.16% on the Day-4 holdout test, outperforming LSTM, XGBoost, and a vanilla Transformer in liquid classification.

Overview

Grasp → sense through the wall → recognize

Opaque bottles hide their contents from cameras; shaking, tilting or opening adds extra actions; and most non-visual approaches need specialized end-effector hardware. Touch2Know turns the grasp itself into the sensor. A multi-finger enveloping grasp stabilizes containers of different diameters and materials, keeps distributed contact around the bottle, and provides several sensing viewpoints at once, so the measurement is less sensitive to small placement errors and matches how a robot really handles a bottle.

Touch2Know overview: UR5-mounted dexterous hand grasping sealed bottles, dataset of 19 liquids and 4 containers, and the sensing-to-classification pipeline
Figure 1. Touch2Know overview. A UR5-mounted dexterous hand performs an enveloping grasp and records 4-channel mutual-capacitance signals through sealed containers using five on-hand flex-PCB electrodes. The dataset covers 19 liquid classes and four container types over four days.
01 · Grasp
Adaptive enveloping grasp
Each finger closes independently until fingertip force sensors report stable contact, with a bounded extra closure to avoid over-compressing soft PET.
02 · Sense
Mutual-capacitance channels
Thumb acts as Tx; index, middle, ring and pinky are four Rx channels. The field path crosses the wall and couples into the liquid.
03 · Calibrate
Per-trial baseline
A 5 s environment baseline recorded before every grasp is subtracted, removing drift from grounding, cabling, humidity and nearby objects.
04 · Classify
CapTF
A CNN–Transformer reads the 4-channel time series and jointly predicts liquid class and container type.
Sensing hardware

Guarded, shielded flex-PCB electrodes on a dexterous hand

Electrodes are fabricated on an 18 µm Cu / 25 µm PI / 18 µm Cu laminate. The object-facing side carries a rectangular sensing pad enclosed by a 1 mm GND guard ring; the back side routes the signal trace inside a grounded copper shield. A coax-like SMA launch and short RF coax runs keep a continuous shielding path to the readout board, and the detachable SMA pair makes any electrode easy to replace. Each electrode is mounted at the second finger joint on 1 mm foam tape, so it sits inside the enveloping contact without interfering with the fingertip 3-axis force sensors.

Flex-PCB electrode workflow: material stack-up, fabrication, electrode structure, SMA coax launch, coax cabling, connector, system integration on the hand, and readout board
Figure 2. Flex-PCB electrode workflow and system integration for multi-channel mutual-capacitance sensing, from material stack-up (a) to on-hand integration (g) and readout (h).
Electrodes
Two sizes to match finger geometry: 0.9 × 2.3 cm (thumb, middle, ring) and 0.9 × 1.5 cm (index, pinky).
Readout
Tashan development board with a 24-bit σ–δ capacitance-to-digital converter, streaming synchronized multi-channel samples logged at 25 Hz.
Hand
BrainCo dexterous hand, 6 actuated joints / 10 DoF, with capacitive 3-D fingertip force sensors (0–25 N, ~0.1 N resolution) used for contact detection.
Grasp controller
Event-driven state machine: precontact → closing → hold → release. Contact detected per finger with hysteresis; hold freezes posture for an 8 s sensing window.
Learning

CapTF: Capacitive Temporal Fusion

A single grasp yields a 4-channel sequence that passes through distinct phases: a short precontact window, a variable-length closing transient, and a long steady-state hold. CapTF is built around that structure. A multi-scale 1-D CNN front-end (kernels 7 → 5 → 3, with a residual block) resolves local contact dynamics; a pre-LN Transformer encoder with a learnable positional encoding captures long-range phase-to-phase context; and mean, max and attention pooling are concatenated into a shared representation.

Because the measured coupling depends jointly on the liquid and the container wall, CapTF is trained as a multi-task model with two heads, one for liquid class and one for container type. Predicting the container encourages the shared backbone to learn features that are less entangled with container-dependent offsets, which improves liquid separability by roughly 3 pp over a single-task variant.

CapTF architecture: 1-D convolutional embedding, Transformer encoder with learnable positional encoding, multi-head pooling, and two classification heads
Figure 3. CapTF architecture. A multi-scale 1-D CNN embeds the 4-channel mutual-capacitance input; a Transformer encoder with learnable positional encoding captures temporal dependencies; multi-head pooling (mean + max + attention) produces a shared representation fed to two task-specific heads.
Results

Fine-grained liquid recognition through sealed bottles

We evaluate on 19 liquids (water, ethanol, grape juice, milk, milkshake, oil, soy sauce, syrup, vinegar, handwash, dishwashing liquid, plus salt water at 2/4/6/8% and sugar water at 5/10/15/20%) in four containers (hard and soft PET, thin and thick glass), under two splits: a stratified 70/30 balanced split and a Day-4 holdout split that trains on Days 1–3 and tests on Day 4. Splits are at the trial-file level to prevent leakage. Delta denotes per-trial baseline-subtracted features; Raw uses the unprocessed signal.

Delta features Raw features
Model · Liquid accuracy Balanced Day-4 Balanced Day-4
CapTF (ours) 95.07% 90.00% 80.15% 68.82%
CapTF-Liq (single-task) 92.11% 86.51% 76.75% 65.79%
LSTM 91.56% 83.68% 78.73% 67.11%
XGBoost 79.39% 77.37% 57.35% 51.84%
Vanilla Transformer 47.15% 45.00% 51.86% 43.29%
Table 1. Liquid classification accuracy (19 classes). CapTF container accuracy: 99.34% (balanced) / 98.16% (Day-4) with delta features; 96.82% / 97.24% with raw features. Macro-F1 tracks accuracy closely for all models; see the paper for the full table.
Per-trial baseline subtraction is the single biggest lever
80.15% → 95.07% balanced · 68.82% → 90.00% Day-4
Raw mutual-capacitance signals carry large trial- and day-dependent offsets from grounding, cable routing and humidity. A 5 s baseline before each grasp removes them and widens inter-class margins; LSTM and XGBoost benefit just as much.
Multi-task learning improves liquid separability
92.11% → 95.07% balanced · 86.51% → 90.00% Day-4
Predicting the container alongside the liquid disentangles wall-dependent coupling from liquid-dependent permittivity and conductivity, consistent with the through-wall sensing physics.
Local + global temporal modelling both matter
Vanilla Transformer: 47.15% · CapTF: 95.07%
Self-attention alone cannot resolve the short contact transients; the CNN front-end supplies local structure while the Transformer links closing dynamics to the steady-state hold.
Remaining errors are the hard, small-margin pairs
Day-4: −5.07 pp vs. balanced
Confusions concentrate among water-based mixtures, adjacent salt/sugar concentrations and visually similar household liquids; liquids with distinctive electrical responses such as oil stay near-perfect across days.
Liquid and container confusion matrices under balanced and Day-4 splits, and a t-SNE visualization of pooled embeddings
Figure 4. CapTF with delta features: (A) liquid confusion matrix, balanced split; (B) liquid confusion matrix, Day-4 holdout; (C, D) container confusion matrices; (E) t-SNE of pooled embeddings for the 19 liquid classes.

Limitations and next steps

The current study uses four bottle types filled near-full (450 mL). Future work will cover broader container geometries, fill-level variation, temperature effects, calibration transfer, open-set recognition, direct prediction of liquid relative permittivity, and integration with downstream manipulation such as pouring and handover.

IROS 2026

Poster

Can't see the poster above? Open the PDF directly.
Citation

BibTeX

@inproceedings{deng2026touch2know,
  title     = {Touch2Know: Dexterous In-Hand Liquid Sensing via
               Mutual-Capacitance Fingerprints},
  author    = {Deng, Ruixiang and Shangguan, Zhegong and Hu, Yang and
               Bai, Haozheng and Li, Tingcheng and Cangelosi, Angelo
               and Yang, Wuqiang},
  booktitle = {IEEE/RSJ International Conference on Intelligent Robots
               and Systems (IROS)},
  year      = {2026}
}