telecosmik-narrative-arguments
Arguments and Reasoning for the TeleCOSMIK Presentation
Date recorded: 2026-10-09.
This record is for discussing the rationale of the presentation: why we do human-to-robot retargeting, why we use COSMIK and MPC, and how these choices connect to the broader goals of robotics research. Numbers are for precise referencing and do not correspond to slide page numbers.
Numbering and Record Format
P Denotes a Proposition
Here, P1 and R1 are internal numbers within this record, not the four-digit research record numbers of the Vault. Each proposition uses a fixed number, e.g. P1. New propositions continue the numbering; they are not renumbered when presentation order or page numbers change. P1 through P4 retain the meanings used in this round of discussion.
Each record entry contains:
- Content: a statement that can be discussed independently.
- Nature: fact or observation, input condition, goal, route choice, or judgment to be verified.
- Discussion status: agreed, adopted wording, to be discussed, or not adopted.
- Basis: user confirmation, specific experience, literature, mathematical reasoning, or experimental record.
"Premise" and "conclusion" are the roles a proposition plays in a particular inference, not two permanently distinct types of node. A proposition can be the conclusion of R1 and also a premise of R2.
R Denotes a Reasoning Step
Each time premises are connected to a conclusion, a fixed number is used, e.g. R1. Each record entry contains:
- Natural-language argument: first write out the premises, connecting reasoning, and conclusion in full, marking the corresponding P numbers alongside the relevant sentences. The reader should not have to flip back to the proposition list to understand.
- Premises: which P numbers are referenced.
- Conclusion: which P number is supported.
- Mode: D, I, or A.
- Rationale: the specific connection between premises and conclusion.
- Scope and open items: how far this step can reach, and what is still missing.
Symbolic shorthand comes after the natural-language argument; it is for tracking and referencing, not a substitute for explanation. P and R references use Obsidian section links, displaying the short number, with hover preview of the corresponding content; these links are not placed inside code blocks.
Reasoning modes:
- D, deduction: under explicit conditions, the conclusion follows from the premises. Check whether the premises hold, and whether the inference is valid.
- I, induction: forming a more general judgment from multiple observations or experiments. Record the sample, comparison conditions, and scope of applicability.
- A, abduction: proposing a candidate explanation or plan with reasons. This record also uses A to mark scheme-level inferences made around research goals; it is not a proof of uniqueness or optimality.
During discussion, use the shorthand:
This means P1 and P4 support our choice of P2; it does not mean P2 necessarily holds. Only D expresses "under these premises, one can conclude"; A or I must not be rewritten as unconditional necessities.
Input conditions that have not been made explicit should be listed separately. Do not hide them behind a single "assume this is so", and do not register hoped-for effects as facts.
Two Parts of the Presentation
S1 is the narrative: what problem we are addressing, why it is worth doing, why we chose this route, and the trade-offs against alternative routes.
S2 is the method: what relationships retargeting preserves, how the mathematical objectives and constraints are formulated, and how MPC generates robot motion.
Internal argumentation may have multiple branches. The presentation selects one easy-to-follow path from them; it does not need to read out every P and R in order. First make the argument coherent, then decide on page breaks.
Propositions
P1 Long-term goal for robot applications
We want to extend robot applications from workstations arranged for specific tasks to environments that people use in daily life; the robot needs to adapt to varying objects, layouts, and operating conditions.
Nature: goal. Discussion status: agreed. Basis: user explicitly confirmed during this round of discussion.
The contrast between an industrial workstation and a kitchen illustrates this extension; it does not imply that dedicated automation is obsolete, nor that all factories are free of variation.
P2 Leveraging human demonstrations to acquire manipulation skills
We want to leverage human demonstrations to help robots acquire manipulation skills, without having to specify the complete robot motion for each case.
Nature: route choice. Discussion status: agreed. Basis: user explicitly confirmed, and pointed out that this step is abductive rather than deductive.
P3 The retargeting task within the chosen data collection route
If we want to convert a human's direct manipulation demonstration into robot execution under a different body and a different layout, we need to solve the conversion from human motion to robot motion.
Nature: conditional task argument. Discussion status: to be discussed. Basis: the discussion in this round about "why not just collect data".
Here we call this conversion retargeting. It does not require that all robot learning must use human-body retargeting, nor does it specify a particular loss function or solver.
P4 Humans already possess relevant manipulation experience
Humans already possess experience with the manipulation tasks we are interested in, such as approaching, grasping, carrying, and placing objects.
Nature: fact and observation. Discussion status: adopted wording. Basis: everyday manipulation and demonstration scenarios raised by the user.
It supports leveraging human experience; it does not directly prove that a particular demonstration collection method can improve robot learning performance.
P5 This time we choose to guide the robot online and record execution
We choose to have the human demonstrate manipulation directly in their own workspace, while the robot executes the corresponding motion in real time in the target workspace, recording the robot-side observations, commands, and execution feedback.
Nature: collection route and data objective. Discussion status: adopted wording; specific data fields to be clarified. Basis: user's description of the online system, source scene, target scene, and purpose of data collection.
The human's action record and the robot's execution record should be described separately. Directly completing a teleoperation task and collecting data for subsequent learning are two uses of this route.
P6 The bodies and workspaces can differ
The human on the source side and the robot on the target side have different body structures; the positions, orientations, and table layouts of objects on both sides can also differ.
Nature: input condition. Discussion status: adopted wording. Basis: the human body and the NERO platform in this system, and the two-sided workspace examples in this round of discussion.
These differences impose requirements on retargeting. The absolute trajectory from the source human cannot be directly used as the manipulation trajectory for the target robot.
P7 Human body observations are not robot motion commands
Human body and object observations describe states in the source scene, but do not directly specify the joint motions the target robot should execute.
Nature: data semantics. Discussion status: adopted wording. Basis: different subjects and representations between human body observations and robot execution commands.
For example, recording that the human reached out and grasped the source bottle does not mean we have obtained the command for how the robot should grasp the target bottle at a different location.
P8 Requirements for the collection interface
We want the collection interface to be low-cost, to reduce equipment worn on the operator's body, and to be suitable for long-duration demonstrations; it should not require body markers, VR headsets, or continuously manipulating a leader arm.
Nature: interface goal. Discussion status: agreed. Basis: device, cost, and ergonomic requirements raised by the user.
Reducing wearable and hand-held equipment does not eliminate the burdens of camera placement, calibration, occlusion handling, computational resources, and maintenance.
P9 User's experience with the leader arm
The user already has a leader arm and finds that operating it for extended periods is neither easy nor comfortable.
Nature: specific usage experience. Discussion status: adopted wording. Basis: personal experience explicitly reported by the user in this round.
This explains one practical source of P8; it cannot be generalized to all leader systems or all operators' experiences.
P10 COSMIK provides existing human body observation capability
RT-COSMIK is a real-time, low-cost, markerless human motion estimation tool developed by CNRS colleagues, using RGB cameras and a human body model to obtain human motion.
Nature: fact, describing tool capability and origin. Discussion status: adopted wording. Basis: RT-COSMIK official page.
It describes the available observation capability, not the accuracy or latency tests of a locally integrated version. Hand and object observations cannot all be attributed to COSMIK's contribution.
P11 Using COSMIK to form the body input interface
We leverage the existing COSMIK human body observation capability to use the human's body movements as robot manipulation input.
Nature: implementation choice. Discussion status: agreed. Basis: user-confirmed COSMIK plus MPC retargeting route.
Low cost and no body markers are the reasons for this choice. The human body estimation contribution belongs to the CNRS colleagues; our research needs to explain how these observations become robot motion.
P12 End-effector targets do not fully specify arm configuration
For a robotic arm with redundant degrees of freedom, specifying only the wrist pose does not uniquely determine the configuration of the entire arm.
Nature: geometric fact. Discussion status: adopted wording. Basis: the same wrist target can correspond to different elbow positions and joint configurations.
It shows that body information can distinguish postures that end-effector targets do not distinguish; it does not prove that preserving this distinction necessarily improves manipulation.
P13 Not only transmitting wrist targets
We want retargeting to consider body configuration and hand–object relationships simultaneously, rather than only passing the wrist pose to the robot.
Nature: research goal. Discussion status: direction adopted; value argument to be completed. Basis: user's desire to leverage human body coordination, and the current discussion on body and manipulation relationships.
It still needs to be clarified: is body configuration demonstration content we want the data to preserve, a means to improve object manipulation, or both? "Moving like a human" should not be directly taken as a sufficient justification.
P14 Relationship preservation goals and execution constraints may conflict
The body or hand–object relationships we wish to preserve may require the robot to exceed joint limits, motion constraints, or enter collision zones, and therefore cannot be exactly realized.
Nature: conditional judgment. Discussion status: adopted wording. Basis: the robot's limited motion range and scene geometry.
This entry discusses the conflict between goals and execution constraints. The conflict between body configuration preservation and object manipulation goals is recorded separately in P23.
P15 Execution requirements for operating near people
We want the robot to operate in collaborative environments where other people are nearby, and to respect necessary safety boundaries.
Nature: application requirement. Discussion status: agreed. Basis: collaborative environment requirements explicitly raised by the user in this round.
This is a requirement, not a completed safety conclusion. Body similarity, using MPC, or configuring collision geometry cannot individually prove personnel safety.
P16 Online motion requires continuous updates and must satisfy execution limits
The robot needs to continuously respond to new observations while considering joint limits, motion constraints, and selected collision geometry.
Nature: online execution requirement. Discussion status: adopted wording. Basis: user-confirmed online teleoperation mode and robot execution conditions.
Online does not mean zero latency. The currently modeled robot and scene constraints also do not automatically cover moving people nearby.
P17 Available capabilities of MPC
MPC can jointly optimize a segment of robot motion for objectives and constraints within a prediction horizon, and update upon receiving new observations.
Nature: method capability. Discussion status: adopted wording. Basis: the predictive motion optimization approach used in this system.
This capability still requires a specific motion model, objectives, constraints, and computational budget. Adopting MPC does not automatically yield a feasible solution, nor does it mean all contact decisions have been included in the optimization.
P18 Choosing MPC to coordinate goals and execution constraints
We choose MPC to coordinate, online, the geometric relationships we wish to preserve with the robot's execution limits.
Nature: method choice. Discussion status: agreed. Basis: user-confirmed current system route.
This explains the method's purpose; it does not prove it is superior to per-frame optimization or learning-based control for all tasks, nor does it mean no loss function or weight design is needed.
P19 Current platform and long-term scope
The current fixed-base dual-arm platform is the validation stage; the long-term direction includes whole-body retargeting. Whole-body MPC itself is also not the final application goal.
Nature: research scope. Discussion status: agreed. Basis: user's explicit description of dual-arm, whole-body retargeting, and long-term robot applications.
The presentation needs to separate current results from long-term directions; it must not describe existing dual-arm experiments as a complete whole-body manipulation capability.
P20 Related work exists for key technologies
Camera-based teleoperation, human body following, interaction relationship preservation, and MPC retargeting all have existing prior work.
Nature: literature fact. Discussion status: adopted wording. Basis: AnyTeleop, H2O, OmniRetarget, X-OP.
Entry points to related literature are in Literature Index. When comparing online versus offline approaches, check human input, reference construction, and deployment execution separately. These citations establish precedent; they do not prove one method is superior to another on a shared task.
P21 Comparing methods through concrete requirements
We should compare methods around the same set of tasks and requirements, stating what our approach gains, what costs it incurs, and how far the evidence supports it.
Nature: comparison principle. Discussion status: agreed. Basis: user's discussion about "why adopt our method" and the different design intents of various methods.
Different methods may serve different goals, or may use different solutions for the same goal. The intent explicitly stated by the authors and the intent we infer should be recorded separately.
P22 Comfort for extended use
The camera body interface is more suitable for long-duration demonstration collection than the leader interface being compared.
Nature: comparative judgment to be verified. Discussion status: to be verified; not treated as an established conclusion. Basis: the goal in P8 and the personal experience in P9.
The comparison devices, tasks, operators, and duration of use need to be specified. P9 alone is not a complete controlled comparison.
P23 Body configuration preservation and object manipulation goals may conflict
The body configuration in a human demonstration serves the human's manipulation in the source scene. When the robot's body or target object layout differs, preserving the demonstrated body configuration may not allow the robot's hands to reach appropriate manipulation positions. To complete the corresponding object manipulation, the robot may need to change its body configuration. Therefore, preserving body configuration and completing object manipulation are two goals that may conflict with each other.
Nature: conditional judgment. Discussion status: agreed. Basis: user's description of body differences, scene differences, and manipulation goals, with explicit agreement to add P23.
Body configuration here does not mean copying human joint angles one-to-one; even if what is preserved is a mapped body geometric relationship, it may still be inconsistent with the hand–object relationship required in the target scene. This conflict does not require joint limits or collision constraints to be active before it arises.
Reasoning
R1 The choice of leveraging human experience
We want robots to perform manipulation in environments people use daily, adapting to varying objects, layouts, and operating conditions (P1). Humans already possess experience with these manipulations (P4). Therefore, we choose to leverage human demonstrations, using humans' existing experience to help robots acquire manipulation skills without specifying complete motions case by case (P2).
Status: agreed as a route choice. Scope: does not conclude that demonstration-based learning is the only approach, nor that our data has already improved downstream learning.
R2 Why the chosen data collection route involves retargeting
In our chosen collection route, the human directly manipulates objects in the source scene, the target robot executes corresponding motions in real time, and records its own observations, commands, and execution feedback (P5). However, the robot and the human have different bodies, and the object layouts on both sides can differ (P6); observing the human's motion does not yet provide the target robot's motion commands (P7). Therefore, to complete this collection route, we need to convert the human's manipulation into executable motion for the target robot. This conversion is what we call the retargeting task here (P3).
Status: connection to be jointly reviewed. Scope: this is a conditional task definition, not a requirement for all data collection or robot learning. It does not conclude that COSMIK, MPC, or any particular loss function must be used.
R3 Why we use COSMIK
We want demonstration collection to be low-cost, to reduce equipment on the operator, and to be suitable for extended use, without requiring body markers, VR, or continuously manipulating a leader arm (P8). The user has already encountered the problem of operating the existing leader being neither easy nor comfortable for extended periods (P9). COSMIK, developed by CNRS colleagues, provides existing low-cost, markerless human motion observation capability (P10). Therefore, we leverage this capability to form the body input interface, allowing the human's body movements to enter the robot retargeting system (P11).
Status: choice agreed. Scope: reduced wearables and hand-held devices are interface characteristics; total cost, calibration effort, and long-duration comfort still need separate evaluation.
R4 Why we also consider body configuration
For a robotic arm with redundant degrees of freedom, even when the wrist reaches the same target pose, the entire arm can still have different configurations (P12). Transmitting only the wrist target does not express which body posture the demonstrator used. This gives us a reason to study the preservation of body relationships, and supports the research direction of "not only transmitting wrist targets" (P13). But this step only shows that there is body information that can be transmitted; which information is worth preserving, and what value it has for manipulation, still needs to be explained.
Status: reasoning not yet sufficient. What needs to be completed is why these distinctions are worth preserving: is it demonstration content, task necessity, or motion preference? P12 only proves there is information left unspecified; it cannot prove that more information is necessarily better.
R5 Why we choose MPC
We want the robot to consider both body configuration and hand–object relationships (P13). However, when the body and layout differ, preserving body configuration and completing object manipulation may require different motions (P23); these goals may also conflict with the robot's execution constraints (P14). The robot also needs to respect execution requirements in environments where other people are nearby (P15), and must continuously respond to new observations while considering joint limits, motion constraints, and selected collision geometry (P16). MPC can jointly optimize motion objectives and constraints within a prediction horizon, and update with new observations (P17). Therefore, we choose MPC: on one hand to coordinate trade-offs between different goals, and on the other to satisfy robot execution constraints (P18).
Status: method choice agreed; its relative advantages still need specific comparison. Scope: safety requirements support constraints and reliable execution, but do not prove MPC automatically guarantees personnel safety. MPC also has costs in online computation, modeling, loss design, and weight selection.
R6 Why we compare through requirements and trade-offs
Camera-based teleoperation, human body following, interaction relationship preservation, and MPC retargeting all have existing related work (P20). Therefore, merely listing that we adopted these technologies does not yet explain why this combination suits our task. We choose to compare routes around concrete tasks and requirements, stating what our choices gain, what costs they incur, and what evidence supports the results (P21).
Status: comparison approach agreed. Scope: a combination of requirements does not automatically constitute innovation; specific mathematical differences, experiments, or practical use value are still needed.
What inductive reasoning can currently support
No I-type reasoning completed with controlled experiments from this system has been registered so far. P9 is a genuine usage experience that can support research motivation; it cannot yet derive the general comparative judgment in P22.
When adding I-type reasoning in the future, state the source of the records or experiments, the comparison conditions, the observed results, and the scope the conclusion covers. Data not yet collected must not be filled in as premises.
Three Standing Discussion Questions
How to capture human motion
Corresponds to P8 through P11, P22, and R3. Compare cameras, body markers, VR, and leader systems in terms of observation capability, equipment burden, cost, and usage conditions.
What from the demonstration to transmit to the robot
Corresponds to P3, P6, P7, P12 through P14, P23, and R2, R4. Discusses wrist targets, body configuration, bimanual relationships, hand–object relationships, and scene differences. This is about problem definition and mathematical objectives, not merely choosing an input device.
How to make the robot execute
Corresponds to P14 through P18, P23, and R5. Compares how online optimization, predictive control, and learning-based tracking handle motion, constraints, training, and computational budgets. The retargeting task, MPC solve, controller tracking, and hardware execution are not the same concept.
COSMIK mainly provides human body observation; retargeting defines what to transmit; MPC is the means for optimizing motion; the controller and hardware are responsible for actual execution. Relationship preservation objectives can enter MPC directly; it is not necessary to build a separate trajectory optimization layer externally for this purpose.
Arguments Not Yet Closed
- P2 to P5: leveraging demonstration-based learning does not necessarily require online teleoperation data collection. Why does this round choose real-time robot following, rather than only recording the human body or generating references offline? We need to compare the uses of robot-executed data and live guidance; this step cannot be skipped.
- P13 and R4: what is the specific value of preserving body configuration? Find a real example of "wrist in position but body motion still unsuitable", distinguishing demonstration content from manipulation effect.
- P3 and R2: further clarify the robot data being collected, and how the human demonstration enters execution. We cannot jump from general data needs to a particular interface.
- P22 and R3: how to evaluate the operational burden of long-term use? First delimit the equipment and tasks, then discuss the comfort comparison.
- P15 and R5: which safety requirements are already covered by models and mechanisms, and which are still only application requirements? People observation, robot collision avoidance, and body similarity should be treated separately.
- P21 and R6: among the current tasks, which requirements are primary and which are acceptable costs? Use this to select related work and comparison conditions.
- S2: before freezing the method narrative, verify against the actual algorithm version. The PI wants to include more manipulation decisions in MPC—this is a development direction and should not substitute for the description of the current implementation.
From Arguments to Narrative
The candidate main line for S1 is P1, P2, P5, P3, then connecting P11, P13, and P18. The interface, relationship preservation, and execution constraints are three branches of reasoning, not a chain of necessary conclusions. Each choice should state which requirement it addresses.
S2 expands P13 and P18: how to express relationships, how to handle conflicts, how to formulate the optimization problem, and how the solve and execution connect.
The left-right comparison confirmed on the second page is preserved: an industrial workstation arranged around a fixed task versus a daily environment designed for people. Related work serves the argument, explaining alternative routes and trade-offs at the corresponding choice points; a detailed comparison table can serve as backup material and does not need to be explained row by row in the main presentation.
Low equipment burden is a system usage value; which relationships to preserve across different bodies and scenes, and how to generate executable motion, is the research question. The two can be presented together but cannot substitute for each other's evidence.