AIfred: Augmented Learning through
Functional Robotic Embodiment at the Desk

Gregorio OrlandoCyPhy Life, IE University, Spain
Milan GroshevCyPhy Life, IE University, Spain
Eduardo Castelló FerrerCyPhy Life, IE University, Spain

Submitted to the IEEE International Conference on Robotics and Automation (ICRA 2027)

Abstract

Desk-based learning and creative activities benefit from handwritten engagement, yet generative AI tools deliver guidance on a separate screen, creating a gap between where users think and where assistance appears. We present AIfred, a desk-based robotic arm with an end-effector-mounted projector that places AI-generated guidance alongside handwritten work. AIfred combines workspace perception, context-aware content generation, and robot-mediated projection to support math assignments, image generation, and drawing. In a user study (n = 36) comparing AIfred with ChatGPT (GPT-5.6 Luna) on a laptop, both systems performed comparably while assistance was available in the math assignment (6.7 vs. 7.3/10, p = .41), but AIfred yielded 60% higher short-term learning transfer once assistance was withdrawn (7.0 vs. 4.4/10, p = .003). Independent art and design professors ranked drawings produced with AIfred best in 33 of 36 cases. These findings indicate that spatially co-located AI assistance benefits tasks whose guidance shares a spatial frame with the work.

01

Why it matters

AIfred artistic overview: a robotic projector assists a person working at a desk
AIfred combines a robotic arm, a projector, and an overhead camera to deliver AI assistance directly on the user's desk.

Many learning and creative activities unfold on paper at a physical desk, where handwriting encourages reformulation rather than passive transcription. Current AI tools, however, live on a screen: users must switch context between the paper they work on and the display where guidance appears.

AIfred closes this gap. It perceives the user's physical task, generates context-dependent support, and projects it back onto the desk, beside the work.

02

Hardware

AIfred is built from off-the-shelf components arranged around an ordinary desk. The projector rides on the robot, so the guidance can move to wherever the user is working.

AIfred hardware components arranged around the workspace
Experimental setup: robotic arm (1), mini projector (2), projection output (3), OptiTrack motion capture (4), trackable object (5), and overhead USB camera (6).
  1. Robotic armPositions and orients the projector over the desk.
  2. Mini projectorMounted at the end-effector; it displays the AI-generated content.
  3. Projection outputThe guidance appears on the desk, right next to the paper.
  4. OptiTrack motion captureTracks the robot base and the trackable object in real time.
  5. Trackable objectThe user moves it to choose where content appears; an inverse-kinematics solver points the projector at it from a fixed height.
  6. Overhead USB cameraGives a top-down view of the task on the desk.
03

Software pipeline

AIfred's spatial assistance pipeline turns a physical task into co-located guidance in three stages.

01

Workspace perception

An overhead camera captures the desk; a pointing gesture selects the region where assistance is needed.

02

Content generation

A multimodal model interprets the selected region under the active interaction mode and generates task-relevant guidance.

03

Robot-mediated projection

The robotic arm projects the guidance onto the desk at the location indicated by a tracked object.

MediaPipe continuously tracks the user's hand. When the user points at part of the page, AIfred captures that region and sends it, together with the active interaction mode, to Google's Gemini models: gemini-3.8-flash for scene understanding and reasoning, and gemini-2.5-flash-image for image generation.

The prompts are structured to favour scaffolding, such as hints, formulas, and analogous examples, over direct answers, so the user still does the thinking. The result is sent to the robot, which moves the projector toward the trackable object and displays the content on the desk. The whole loop runs without a screen, keyboard, or mouse: the user points, reads, and keeps working on paper.

04

Interaction modes

Solving exercises, sketching ideas, and learning to draw are still done by hand on paper. AIfred supports each of them with a dedicated mode, and every interaction follows the same three phases: the user shows the work, the robot projects support, and the user completes the task.

AIfred's homework, image-generation, and drawing interaction modes
AIfred interaction scenarios by mode (math homework, generate image, and draw) and phase: Phase 1, user interaction; Phase 2, robot projection; Phase 3, result.

Math homework

What you do: write the problem on paper, for example a quadratic equation, and point at it.

How AIfred helps: it projects step-by-step support beside the exercise, such as the relevant formula and a simpler worked example, without giving away the answer. You solve the problem yourself, by hand.

Generate image

What you do: sketch an idea on paper and add a few words describing it, like "rocket car in space".

How AIfred helps: it turns the sketch into a finished digital image and projects it next to your drawing, taking your idea from paper to a polished visual.

Draw

What you do: make a first attempt at a drawing and point at it.

How AIfred helps: it projects a drawing tutorial for your subject onto the desk, so you can follow the guide stroke by stroke and produce an improved drawing on your own paper.

The trackable object doubles as a controller: move it and the arm follows, rotate it to change page, and lift it to switch mode.

05

Results

We ran a mixed-design user study with 36 participants from across the university. Half used AIfred and half used ChatGPT (GPT-5.6 Luna) on a laptop. Everyone solved a quadratic equation, turned a sketch into a digital image, and drew a figure three times: without help, with their assigned system, and with the other system. About 35 minutes later, they solved a second equation with no assistance at all, to measure what they had actually learned.

60%

higher short-term learning transfer

With assistance withdrawn, math scores were 7.0/10 after AIfred vs. 4.4/10 after ChatGPT (p = .003).

33/36

drawings ranked best

Three independent art and design professors ranked AIfred-assisted drawings first in 92% of cases (Kendall's W = .86).

98%

fewer context switches

On average, 1 physical-digital switch with AIfred vs. 63 with ChatGPT, with no loss in perceived productivity (p = .61).

Math assignment and short-term learning transfer

Scatter plot of math grade with assistance against math grade without assistance for AIfred and ChatGPT participants
Final grade with ChatGPT or AIfred assistance versus the grade on a later problem with no assistance.

Each participant solved a quadratic equation, picked from a pool of five of comparable difficulty, with help from their assigned system. About 35 minutes later, after the image and drawing tasks, they solved a second equation with no help at all. Solutions were graded blind by four independent AI grading agents (Gemini Flash 3.6) using the same eight-step rubric, from problem setup and method choice through the discriminant, both roots, simplification, and verification. The agents agreed closely, differing by only 0.53 points on average.

In the chart, each small dot is one participant: their grade with assistance on the horizontal axis, and without it on the vertical. Dots on the diagonal kept the same grade; dots below it lost ground once the help was gone. The large markers show each group's average.

While assistance was available, the two groups scored about the same (6.7/10 with AIfred vs. 7.3/10 with ChatGPT, p = .41). The difference appeared once the help was taken away. The ChatGPT group fell to 4.4/10, a 40% drop, while the AIfred group held steady at 7.0/10, even rising by 4.5%. AIfred users scored 2.6 points higher, a 60% higher short-term learning transfer (p = .003).

We think the reason is how the guidance is used. An answer on a screen can be copied onto paper with little thought. AIfred projects hints, formulas, and analogous examples instead of solutions, so users have to reinterpret them and write out every step themselves. That extra effort is what stays with them when the help is gone.

Drawing quality

Stacked bar chart of drawing ranks for baseline, ChatGPT, and AIfred conditions
Rank of each participant's three drawings, from best (1st) to worst (3rd).

Each participant drew three times: without help, with ChatGPT, and with AIfred. Three independent art and design professors ranked every set, and largely agreed (Kendall's W = .86). Drawings made with AIfred were ranked first in 33 of 36 cases and never last. ChatGPT drawings were mostly second (28 of 36), and unassisted drawings were last in 31. AIfred drawings improved both overall form and finer details such as proportions and line quality. Drawing is where co-location matters most: the reference sits directly on the paper, so there is nothing to re-align after looking away.

Physical-digital context switching

Scatter plot of task completion time against number of context switches for AIfred and ChatGPT
Observed physical-digital context switches and task completion times. Faint points are individual trials; outlined markers are condition means.

We counted every time a participant's attention moved between the desk and the screen. ChatGPT users switched 63 times per task on average; AIfred users switched once, a 98% reduction. Yet when asked, both groups rated context switching as barely disruptive (1.3/5 with AIfred vs. 1.7/5 with ChatGPT). People are so used to jumping between paper and screen that they no longer notice the cost, even though it coincided with lower learning transfer and weaker drawings. AIfred tasks took longer (350 s vs. 260 s), but that time went into working on paper, not into looking away from it.

Before and after

Two examples from the study. Drag each divider to compare what the participant put on the desk with the result obtained with AIfred.

Draw mode

A first unassisted attempt at an elephant, then the same drawing after following AIfred's projected tutorial.

Drawing task inputINPUT
Drawing task output with AIfredAIFRED

Generate image mode

A pencil sketch of a "rocket car in space", then the digital image AIfred generated from it and projected back onto the desk.

Image-generation task inputINPUT
Image-generation task output with AIfredAIFRED

How it felt to use

What participants reported matches what they did. AIfred users rated perceived learning support far higher (4.5/5 vs. 3.4/5, p < .001), and they were the ones who kept their math skills once help was withdrawn. They also found it more innovative (4.5 vs. 3.5) and more satisfying (4.2 vs. 3.5).

AIfred asked more of them: cognitive demand was higher (3.4 vs. 2.9, p = .03) and tasks took longer. Screen-based answers are easy to copy with little thought, while projected guidance has to be reinterpreted and reproduced by hand. That extra effort did not feel like a cost, since perceived productivity was the same in both groups (4.1 vs. 4.0, p = .61).

The biggest gap between feeling and behaviour is context switching. ChatGPT users looked away from their paper 63 times per task, yet rated the disruption as low as AIfred users did. Attention fragmentation works below awareness: people do not notice it, but it still shows up in their learning and drawings.

Takeaway: AIfred makes people work better, not just faster. With guidance on the desk, users stay on the paper, think harder, and keep more of what they practise. The benefit is largest when the guidance and the work share the same space, as in drawing. Robot-mediated projection is not a general replacement for the screen, but a targeted tool for hands-on learning and making.