Deep dive / Explainers

When Vision Is Not Enough: What Touch Adds to Robot Manipulation

Cameras describe a scene before contact. Tactile sensors reveal what happens at the hidden interface where grasp stability, slip, force and material response are decided.

When Vision Is Not Enough: What Touch Adds to Robot Manipulation main imageAI-generated image
The AI-generated editorial illustration explains tactile sensing and does not represent measured data.

AI-generated editorial illustration. It explains tactile sensing at a robot gripper and does not represent measured data.

1. The central question

A camera can show a robot where a cup is, estimate its shape and guide a gripper toward it. The most consequential moment then disappears: once the fingers close, the contact surfaces are hidden. Vision alone cannot directly tell whether pressure is concentrated on an edge, whether the object has begun to rotate inside the grasp, or whether a soft package is being crushed.

What does touch add? It provides information generated by physical interaction rather than distant observation. Depending on the sensor, that information can include contact location, surface geometry, normal and shear deformation, aggregate force, vibration and incipient slip. Touch does not replace vision. It closes the information gap between seeing an object and controlling what happens when robot and object become mechanically coupled.

2. Conceptual foundation

Three terms matter. Normal force acts perpendicular to a contact surface; shear force acts along it. Slip begins when friction is insufficient to maintain the existing contact. A robot does not always need each quantity in calibrated physical units. A policy may only need a reliable pattern indicating that contact has shifted or that grip should tighten.

Tactile sensing also differs by spatial scale. A wrist force/torque sensor measures the combined load transmitted through an arm. It is useful for detecting collision or regulating insertion force, but cannot easily separate several fingertip contacts. A tactile fingertip provides local information at the contact patch. A skin distributed over a hand, arm or body can detect contacts across a larger area, usually at lower spatial resolution or with more wiring and calibration complexity.

3. How the system works

A tactile control loop begins when a compliant sensor surface touches an object. The contact changes an electrical, magnetic or optical signal. Signal processing converts that raw response into a representation: a pressure map, a deformation image, marker displacement, estimated load or learned feature vector. The robot combines this representation with joint state and camera observations, then changes grip force, pose or trajectory. Because contact evolves quickly, sensing, synchronization and control latency matter as much as static accuracy.

Vision-based tactile sensors place a camera behind a soft elastomer. Illumination reveals how the surface deforms against an object; markers can expose lateral motion associated with shear and slip. Magnetic skins embed or attach magnetic material to a soft interface and measure how the field changes as the surface deforms. Force/torque sensors use strain-sensitive structures to measure aggregate loads. Each converts mechanics into data differently, so their outputs are not interchangeable without a learned or engineered mapping.

4. Major approaches

Wrist and joint force sensing. These sensors observe the net mechanical consequence of contact. They are effective for compliant motion, collision detection and tasks such as insertion where the direction of resistance matters. Their weakness is localization: a single reading may combine object weight, acceleration and multiple contacts.

Vision-based tactile fingertips. GelSight measures the deformation of a soft reflective surface with an internal camera. Its high-resolution contact geometry can support shape, texture, force-related and slip inference. DIGIT packages a related optical principle into a compact form intended for multi-finger hands. This family delivers information-rich images but needs a camera, controlled illumination, optical packaging and a replaceable elastomer surface.

Magnetic tactile skins. ReSkin separates sensing electronics from a replaceable magnetic interface. Deformation changes the magnetic field measured beneath the surface. AnySkin extends this direction with an adhesive-free interface designed to be replaced rapidly while preserving learned behavior across instances. These sensors emphasize durability, serviceability and data reuse rather than camera-like spatial imagery.

Multimodal fusion. A policy may combine vision, proprioception and touch at several stages. Early fusion concatenates synchronized sensor features; later fusion lets separate encoders specialize before their outputs influence a policy. A hierarchical controller may use vision for approach and tactile feedback for contact. The appropriate design depends on whether touch corrects occasional ambiguity or drives the whole contact phase.

5. Evidence and examples

The 2017 GelSight paper describes a soft elastomer whose vertical and lateral deformation captures high-resolution contact geometry. Marker motion can support inference of local force, shear and slip, and the paper reports applications in material perception and in-hand localization. The evidence established that a robot fingertip could obtain far more than a binary contact signal. It did not establish a universal calibration across every sensor construction.

DIGIT addresses packaging. Its authors designed a compact, repeatable optical tactile sensor suitable for multi-finger robot hands and demonstrated learned controllers manipulating glass marbles in-hand. That experiment links dense tactile images to a closed-loop behavior on a specific hand. It remains a selected in-hand task, not evidence that optical touch alone solves general dexterity.

ReSkin focuses on replacement and scaling. Soft interfaces wear, and a sensor whose electronics must be discarded or recalibrated after every replacement is difficult to deploy in sustained data collection. ReSkin’s magnetic architecture and learning methods were designed to tolerate fabrication and time variation. AnySkin further decouples electronics from an adhesive-free interface. Its authors report slip-detection experiments and zero-shot transfer of manipulation policies from one sensor instance to another under their tested conditions.

These cases are not a leaderboard. GelSight and DIGIT emphasize rich local geometry; ReSkin and AnySkin emphasize replaceability and reuse. Their tasks, sensor footprints and outputs differ. Together they show that tactile-system design is inseparable from the maintenance and data pipeline around the sensor.

6. Trade-offs

The clearest trade-off is resolution versus durability. A detailed optical imprint can reveal fine contact geometry, yet the exposed soft surface still experiences abrasion and contamination. A simpler magnetic interface may be easier to replace, but provides a different spatial signal and relies on models to interpret field changes.

There is also calibration versus interchangeability. Precise force estimates can require a sensor-specific mapping between deformation and load. Replacement changes materials, thickness and assembly. Cross-instance learning aims to make policies tolerant to those changes, but tolerance can reduce sensitivity to subtle signals. Serviceable hardware is valuable only when the data distribution remains usable after maintenance.

Bandwidth versus reaction time is another constraint. High-resolution tactile images create large streams that must be synchronized with cameras and robot state. Compressing them reduces compute but may discard a brief slip event. A local reflex can react quickly, while a large multimodal policy may understand more context but respond more slowly.

Finally, touch has a coverage versus complexity problem. Two fingertip sensors are manageable but only sense where the fingertips touch. Full-hand or body skin increases coverage, cables, electronics, calibration surfaces and failure points. The right sensor placement follows the task, not the aspiration to make every robot human-like.

7. Current limitations

Tactile datasets are small and heterogeneous compared with vision datasets. Sensors vary in shape, material, sampling rate and output. Even sensors built from the same design can respond differently after fabrication, installation or wear. This makes cross-laboratory reuse difficult and complicates training of general models.

Ground truth is also hard to obtain. Contact pressure, shear and geometry occur together, while calibration fixtures isolate them imperfectly. Learned estimates can be accurate inside the calibration range and unreliable on new materials or contact shapes. Soft surfaces age; optical windows become dirty; magnetic interfaces can shift. Published demonstrations rarely report long maintenance cycles or policy behavior after repeated sensor replacement.

8. Open questions

Can a common tactile representation span optical images, magnetic fields and force arrays without erasing the information unique to each sensor? A useful answer would allow shared pretraining while preserving fast, hardware-specific control. Researchers also need benchmarks that test slip recovery, gentle handling, insertion and sensing after wear—not only contact classification.

Another question is architectural: should high-frequency tactile reflexes remain in small local controllers while slower multimodal models plan the task? That division could protect latency and safety, but may prevent an end-to-end policy from learning subtle coordination. The field also needs evidence about the minimum tactile coverage required for useful humanoid work.

9. Editorial synthesis

Touch is most credible when tied to a specific information failure. Use wrist sensing when aggregate load and compliance matter. Use high-resolution optical fingertips when local geometry and object motion inside a grasp are central. Use replaceable magnetic skins when repeated contact, maintenance and cross-sensor data reuse dominate the engineering problem.

The larger lesson is that sensing quality cannot be reduced to resolution. A sensor must survive, be replaced, synchronize with control and produce signals a policy can reuse. AnySkin’s attention to cross-instance transfer is important because it treats maintenance as part of learning. GelSight and DIGIT remain important because they show how much hidden contact structure is available. Progress will come from combining those virtues—rich signals, replaceable surfaces and fast control—rather than declaring one modality sufficient.

10. Key takeaways

  • Touch reveals hidden contact conditions that external cameras cannot directly observe.
  • Wrist force sensing, optical fingertips and magnetic skins answer different mechanical questions.
  • High spatial resolution improves contact understanding but increases packaging, bandwidth and maintenance demands.
  • Replaceability matters only if learned policies remain usable across sensor instances and wear states.
  • Vision and touch are complementary: vision guides approach, while touch stabilizes and corrects contact.

11. Further reading

  1. What Is Robot Perception? — understand how cameras, proprioception and touch form a shared state estimate.
  2. How Dexterous Robot Hands Work — connect tactile sensing to actuators, tendons and grasp control.
  3. Why Robots Generate Actions in Chunks — examine how sensory feedback changes predicted motion sequences.

12. Sources & evidence

Related reading

Sources

Accessed: 2026-09-02

  1. AnySkin: Plug-and-play Skin Sensing for Robotic Touch
    Raunaq Bhirangi, Venkatesh Pattabiraman, Enes Erciyes, Yifeng Cao, Tess Hellebrekers and Lerrel Pinto. Submitted 2024-09-12.
    https://arxiv.org/abs/2409.08276
    https://any-skin.github.io/
    Supports: replaceable adhesive-free interface, slip detection, policy learning and cross-instance transfer.

  2. DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation
    Mike Lambeta et al. IEEE Robotics and Automation Letters, 2020.
    https://arxiv.org/abs/2005.14679
    Supports: compact optical tactile hardware and learned in-hand manipulation evidence.

  3. GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force
    Wenzhen Yuan, Siyuan Dong and Edward H. Adelson. Sensors, 2017.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC5751610/
    Supports: optical tactile sensing principles, contact geometry, local force, shear and slip inference.

  4. ReSkin: Versatile, Replaceable, Lasting Tactile Skins
    Raunaq Bhirangi, Tess Hellebrekers, Carmel Majidi and Abhinav Gupta. CoRL 2021, PMLR 164.
    https://arxiv.org/abs/2111.00071
    https://proceedings.mlr.press/v164/bhirangi22a.html
    Supports: magnetic tactile sensing, replaceable interfaces and cross-instance robustness motivation.

  5. GelSight Wedge: Measuring High-Resolution 3D Contact Geometry with a Compact Robot Finger
    Rui Li et al. ICRA 2021.
    https://arxiv.org/abs/2106.08851
    Supports: compact fingertip packaging and contact geometry under visual occlusion.