TailorCoPilot was associated with a higher chance of task success among novices than the skill-appropriate baseline tools used in a comparative evaluation. Across the main test, the odds of success were 2.31 times as high with TailorCoPilot as with baseline, with a 95% confidence interval from 1.42 to 3.86 and a reported p-value of .001.
The advantage varied by experience and task difficulty. Novices showed an advantage at every difficulty level. Advanced novices were closer to baseline on simple tasks, but the separation was wider on medium and complex tasks.
A test built around repeated revisions
The main evaluation involved 20 participants, split between 10 novices and 10 advanced novices. Each person completed 12 trials over two sessions, with six trials per session. Each session included three simple tasks, two medium tasks and one complex task. Workflow order was counterbalanced, and task assignments were randomized and balanced without repeating a task for the same participant.
Researchers compared TailorCoPilot with skill-appropriate baseline tools. Its trace layer treated every commit as a viable garment-pattern state. Each state combined structured 2D pattern geometry, 3D draped geometry and multimodal metadata, and was stored only after validation alongside the semantic sequence of operations that produced it.
The system was trained on 93 traces. Of those, 56 were textbook traces covering three to 10 states, while 37 came from live production and covered 10 to 30 states. A trace-level split separated training from validation, and approximately 2,500 translation pairs were used to fine-tune Qwen3-VL-8B.
Speed came with a lighter workload
The study also included a formative phase with 13 participants: six novices, four advanced novices and three experts. Sessions lasted about 60 minutes, divided between a 30-minute interview and a 30-minute task session. The experts had more than five years of industry experience.
Researchers used retrospective think-aloud protocols and screen recordings in the formative work. They analyzed transcripts, task recordings and artifacts with reflexive thematic analysis, then had two researchers review the material and refine the themes.
TailorCoPilot users also completed tasks in less time than baseline users. The reported p-value for the overall time comparison was below .001, and the time difference varied with task difficulty, with p = .018. The assistant was generally faster on medium and complex tasks, although CAD baseline performance remained competitive for some simple tasks completed by advanced novices.
NASA-TLX, the workload measure used by the study, was lower with TailorCoPilot, with a reported p-value below .001. Exit interviews indicated that participants retained a sense of agency and ownership during revision.
Experts saw a difference in the finished patterns
For artifact quality, five blinded experts rated a stratified random sample, with two independent ratings for each sampled artifact. The ratings showed strong inter-rater reliability, with an ICC of 0.81, and overall artifact quality was higher with TailorCoPilot, with the clearest gains in structural balance and conformity to the reference and smaller but consistent gains in sewability.
On the study's 1-to-5 rating scale, TailorCoPilot exceeded baseline on every reported dimension in both expertise groups. Among novices, the overall-quality mean was 3.92, with a standard deviation of 0.82, compared with a mean of 1.62 and a standard deviation of 0.70 for baseline. Among advanced novices, the corresponding mean was 4.27, with a standard deviation of 0.71, versus 2.87 and 0.95 with baseline.
Taken together, the comparison points to a consistent advantage for TailorCoPilot in the tested workflow: higher task-success odds, shorter completion times, lower reported workload and stronger expert-rated artifacts. The pattern was most pronounced on harder tasks among advanced novices, while novices showed an advantage across the full difficulty range.
Paper data and sources
Original title: TailorCoPilot: Enabling Agentic Pattern Making with Version-Controlled State Tracking
Authors: Yuexin Sun, Zhaohui Wang, Ruiyang Liu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text