top of page

The Right Data for Robots: Ego-Centric vs Stereo Camera Rigs for Robotics, with or without Tactile Gloves

Aug 31
9 min read

A robot can watch a hand pick up a cup and still miss the most important part: how the fingers know when to stop squeezing. Vision tells a robot where objects are. Touch tells it what contact feels like. For robotics data collection, the best rig often needs both.


Video data collection rigs sit at the center of modern robot learning. They help capture how people move, grasp, assemble, sort, inspect, and navigate. That data can train systems for imitation learning, teleoperation, manipulation, quality inspection, and human-robot interaction.


Two common approaches stand out: ego-centric camera rigs and stereo camera rigs. Both can record useful video, but they answer different questions. Ego-centric systems show the task from the human operator’s point of view. Stereo systems capture depth and spatial geometry from two coordinated cameras. When pressure-sensing tactile gloves enter the setup, the dataset becomes far richer because it links what the operator sees with what the hand feels.


Eye-level view of a robotic data collection setup with wearable cameras and tactile gloves
Vision and touch together create a fuller record of human manipulation.

Why video rig choice matters in robotics data collection


Robots struggle with tasks that humans do without thinking. Opening a jar, folding cloth, inserting a cable, or packing fruit all require visual judgement and contact control. A camera rig decides what the system can learn from these actions.


A weak rig may capture a good-looking video but miss the fine details that matter. The hand may block the object. The scene may lack depth. The force applied by the fingers may be unknown. Lighting may change. Motion blur may hide an important grip adjustment.


A strong rig captures the task in a way that connects three layers:


  • What the operator sees

  • Where the objects and hands are in space

  • How contact changes during the task


Ego-centric and stereo rigs approach these layers differently.


Ego-centric camera rigs record the human task from the operator’s view


An ego-centric rig places one or more cameras on the person performing the task. The camera may sit on the head, chest, shoulder, wrist, or even near the hand. The goal is simple: record the world as the operator experiences it.


This approach is useful for tasks where attention matters. If a technician looks at a screw hole before inserting a part, the camera captures that viewpoint. If a worker reaches into a bin and selects one item from many, the system records the scene from the same angle the worker used to decide.


Advantages of ego-centric rigs


They capture intent well


An ego-centric view often shows what the person is focused on. In assembly tasks, the camera follows the operator’s gaze and body motion. This can help robot learning systems connect visual attention with action.


For example, in a warehouse picking task, an ego-centric camera can record how a person scans a shelf, identifies a label, reaches toward the correct bin, and grasps the item. A fixed camera might see the body from outside, but it may not reveal the operator’s visual focus as clearly.


They are portable


Ego-centric rigs suit field data collection. They can be used in factories, farms, kitchens, healthcare training environments, retail back rooms, or warehouses without rebuilding the space around cameras.


This is valuable when the same task must be recorded across the USA, Europe, and India, where workspace layouts, lighting, and tools may vary.


They work well for human demonstration data


Many robot learning pipelines rely on demonstrations. A person performs a task, the system records the sequence, and the robot later imitates pieces of that behavior. Ego-centric footage makes the demonstration feel natural because the operator can move freely.


Disadvantages of ego-centric rigs


Hands and tools can block the view


The biggest weakness is occlusion. During close manipulation, the operator’s hands often cover the exact contact area the robot needs to understand. When tightening a small connector, the camera may see only knuckles and tool handles.


Depth can be unclear


A single ego-centric camera does not directly measure distance. A model can estimate depth from video, but estimates may fail when surfaces are smooth, transparent, shiny, or textureless.


Motion can be noisy


Head-mounted cameras move with the body. Fast turns, bending, walking, or reaching can create blur and unstable footage. Good mounting, lighting, camera placement, and synchronization help, but they do not remove the problem entirely.


Close-up view of pressure-sensing tactile gloves gripping a small plastic part
Tactile sensing adds the missing contact layer to visual demonstrations.

Stereo camera rigs add depth and spatial structure


A stereo camera rig uses two cameras separated by a known distance. By comparing the two images, the system can estimate depth in the scene. This resembles how human vision uses two eyes to judge distance.


Stereo rigs may be fixed in the environment, mounted on a robot, attached to a mobile platform, or placed above a workspace. They are especially helpful when the task requires accurate 3D understanding.


Advantages of stereo camera rigs


They provide depth information


Stereo vision helps estimate where objects sit in 3D space. This is useful for robot arms, mobile robots, bin picking, and inspection systems.


For example, in a parts bin, many metal brackets may overlap. A stereo rig can help determine which bracket is on top, how it is oriented, and whether a gripper has enough space to approach.


They support object tracking and measurement


Stereo rigs can track hand position, object pose, and relative motion across frames. When calibrated well, they provide structured spatial data that a robot can use for planning.


In agriculture, a stereo rig mounted near a robotic end effector can help estimate fruit size, fruit position, and branch distance. In logistics, it can help a robot judge parcel dimensions before grasping.


They are more stable when fixed


When mounted on a tripod, frame, robot cell, or ceiling rail, stereo cameras can produce stable footage with less motion shake than wearable cameras.


Disadvantages of stereo camera rigs


They can miss the operator’s point of view


A stereo rig may capture excellent geometry but lack human intent. It can show where the hand moved, but not always what the person looked at or why they chose a specific object.


Setup and calibration take care


Stereo rigs require careful alignment, synchronization, and calibration. Camera spacing, lens choice, lighting, frame rate, and mounting all affect quality. If calibration drifts, depth quality drops.


They may not fit every environment


Fixed rigs work well in controlled workcells, but they are harder to use in cramped, changing, or outdoor settings. A field technician repairing equipment may not have a place to mount a stereo system at the right angle.


Ego-centric and stereo rigs solve different problems


Neither approach wins in every case. The choice depends on the task, the environment, and the type of model being trained.


Ego-centric rigs


Best when the dataset needs the operator’s viewpoint, attention, and natural motion. Useful for wearable demonstrations, mobile tasks, repair work, tool use, and human-centered workflows.

Main strength


Captures the task as the human sees it.

Main weakness


Can suffer from occlusion, shake, and unclear depth.

Typical mounting


Head, chest, wrist, shoulder, or hand.

Stereo rigs


Best when the dataset needs scene geometry, object depth, stable tracking, and accurate spatial relationships. Useful for fixed workcells, bin picking, robot arms, inspection, and navigation.

Main strength


Captures distance and 3D layout.

Main weakness


Can miss human attention and requires careful setup.

Typical mounting


Tripod, frame, robot arm, mobile base, or overhead rail.


For many robotics teams, the strongest answer is a hybrid setup. An ego-centric camera records the operator’s view. A stereo rig records the workspace in 3D. The tactile glove records contact. Together, they create a dataset that is much closer to the real human skill being demonstrated.


Wide-angle view of a stereo camera rig observing a robotic arm and scattered parts
Stereo cameras help robots understand depth, pose, and object layout.

Tactile gloves fill the gap between seeing and touching


Vision alone cannot tell the whole story of manipulation. A video may show fingers touching a sponge, but it may not reveal whether the operator pressed lightly or firmly. It may show a cable insertion, but not when resistance increased. It may show a fruit being picked, but not whether the grip was gentle enough to avoid bruising.


Pressure-sensing tactile gloves capture this missing layer. They can measure pressure patterns across fingers, fingertips, and the palm. Depending on the glove design, they may also include bend sensors, inertial sensors, or time-synchronized markers.


What tactile gloves add to the dataset


Contact timing


The system can detect when the hand first touches an object, when grip force rises, and when the object is released.


Grip strategy


A glove can help identify whether the operator used a pinch grip, power grip, side grip, or palm support.


Force distribution


Pressure data shows which fingers carried the load. This matters for delicate tasks such as handling glassware, electronics, soft produce, or medical training tools.


Error signals


If a person slips, readjusts, or presses too hard, the glove can capture that event. These “almost failed” moments are often valuable for training safer robot behavior.


Why tactile data works well with ego-centric video


Ego-centric video shows the operator’s view. Tactile gloves explain what the hand is doing when the view is blocked. If fingers hide a connector during insertion, pressure data can still show when contact begins and how the grip changes.


This pairing is useful for:


  • Tool handling

  • Assembly and repair

  • Surgical training simulation

  • Cooking and food handling research

  • Assistive robotics

  • Human hand skill analysis


Why tactile data works well with stereo video


Stereo cameras provide 3D context. Tactile gloves provide contact context. Together, they help connect object pose, hand pose, and applied pressure.


For example, a stereo rig can track a gripper-like hand path toward a valve handle. The glove can show how much pressure the operator applied while turning it. A robot can learn not only the path but also the contact behavior.


Real-world examples of rig selection


Warehouse picking


A wearable ego-centric camera works well when recording how people identify items on shelves, read labels, and reach into bins. A stereo rig helps when the goal is to teach a robot arm to estimate item position and plan a grasp. Tactile gloves add value by showing how people adjust grip when lifting soft bags, cardboard boxes, or slippery packaging.


Manufacturing assembly


For cable routing, screw fastening, clip insertion, and part alignment, a hybrid rig is often best. Ego-centric video captures the worker’s view of small features. Stereo cameras capture the 3D layout of parts and tools. Tactile gloves show insertion force patterns and grip changes.


This is especially useful for tasks where success depends on feel, such as snapping a connector into place.


Agriculture and food handling


Robots that harvest fruit or sort produce must handle objects gently. Stereo cameras help estimate size, position, and distance. Ego-centric video can record how skilled workers choose ripe items or avoid branches. Tactile gloves show pressure limits and grip style, which can help reduce damage.


Healthcare and assistive robotics


In rehabilitation research, tactile gloves can record how patients or therapists interact with objects during hand exercises. Ego-centric video provides task context, while stereo rigs measure movement and spatial accuracy. The data can support robot-assisted training tools and adaptive assistive devices. Any healthcare use should follow local safety, privacy, and compliance rules.


Home and service robots


Home tasks are messy and varied. Opening drawers, folding towels, pouring grains, and loading dishwashers involve clutter, occlusion, and many object types. Ego-centric rigs are useful because they capture natural human demonstrations in real rooms. Stereo rigs help when building robot-ready 3D models of the scene. Tactile gloves reveal the contact skills hidden inside each action.


Over-the-shoulder view of a person demonstrating object handling for a robot
Hybrid rigs can capture viewpoint, depth, and pressure in one recording session.

How to choose the right data collection rig


A good rig starts with the task, not the camera catalog. The best setup depends on what the robot must learn.


Use an ego-centric rig when:


  • The operator’s viewpoint matters

  • The task happens across different locations

  • The person must move freely

  • Human attention and sequence are key

  • Setup time must stay low


Use a stereo rig when:


  • Depth is central to the task

  • Object pose and distance matter

  • The workspace can support fixed mounting

  • You need stable 3D tracking

  • Calibration can be managed carefully


Use pressure-sensing tactile gloves when:


  • Grip force matters

  • Objects are soft, fragile, slippery, or deformable

  • The task involves insertion, twisting, pressing, or sliding

  • Hands block the camera during key moments

  • The robot must learn contact-rich behavior


For complex manipulation, combine all three. The data becomes easier to interpret because each sensor covers another sensor’s blind spot.


Data quality matters as much as hardware


A powerful rig can still produce poor training data if the collection process is weak. Teams should plan for synchronization, labeling, calibration, storage, and repeatability.


Key practices include:


  • Time-sync video and tactile signals

  • Record calibration sequences at the start of sessions

  • Use consistent object names and task labels

  • Capture successful and failed attempts

  • Include variation in lighting, object placement, tools, and operators

  • Track metadata such as camera position, glove size, sensor range, and task conditions

  • Protect personal and sensitive data, especially in healthcare, home, or workplace settings


The goal is not just to collect more video. The goal is to collect data that a robotics team can trust.


How Xelec can assist with robotics data collection rigs


Xelec can help teams plan, build, and support video data collection systems for robotics, including ego-centric rigs, stereo camera setups, and pressure-sensing tactile glove integration.


Support may include:


  • Selecting the right camera layout for the task

  • Designing wearable and fixed rig configurations

  • Integrating tactile gloves with video capture

  • Planning synchronization between sensors

  • Building repeatable data collection workflows

  • Supporting deployments across India, the USA, and Europe

  • Helping teams collect richer data for robot learning, manipulation, inspection, and automation research


The right rig can reduce blind spots in the dataset and make robot training more practical. Ego-centric cameras capture the human view. Stereo cameras capture spatial structure. Tactile gloves capture the feel of contact. Used together, they give robotics teams a clearer picture of how skilled actions really happen.


For inquiries or project discussions, contact Xelec at gulshan@xelec.in.



 
 
 

Comments


bottom of page