When we started building Atorie, we made most of our product decisions from instinct: Redouane's eight years as a working stylist, Mia's work on recommendation interfaces, Stefan's understanding of matching systems. That instinct got us to a first version we were genuinely proud of. But instinct has limits. So in late spring 2026 we ran a formal four-week pilot with 40 participants, and we are still absorbing what they told us.
The participants were recruited through an interest form we had linked from early press and from a handful of independent boutique newsletters. We specifically did not recruit from our own extended networks, because people who already like you tend to tell you what you want to hear. We wanted people with strong opinions about dressing, a tolerance for unfinished product, and no particular loyalty to us. We got those people. The feedback was direct.
Who was in the pilot
The 40 participants spread across Los Angeles, New York, Chicago, and Portland. About a third described their relationship to fashion as professional: working stylists, boutique buyers, a few fashion writers at small publications. The remaining two-thirds were what we might call committed amateurs: people who do not work in fashion but spend real time and money thinking about how they dress, who follow independent labels on social media, and who expressed frustration with mainstream discovery tools.
We asked each participant to use Atorie as their primary discovery tool for four weeks. They could save outfits, skip suggestions, and update their style profile during the period. We sent a short check-in survey at the end of week two and a longer exit survey at the end of week four. We also conducted thirty-minute video interviews with 12 of the participants after the pilot ended.
One constraint worth naming: this was not a blind study, participants knew they were testing a product in active development, and that context probably softened some of the criticism. We account for that when we interpret the results, which means assuming that problems people did mention were real problems, and that the absence of a complaint is not the same as an absence of the underlying issue.
What worked better than expected
The outfit context mattered more than the label discovery. We had designed Atorie around the idea that finding new labels was the central value, and we expected feedback to focus on whether the labels were relevant. Instead, the most common positive theme in exit surveys was about seeing garments worn together in context. Participants described understanding how to actually use a piece they had just discovered, because they could see it alongside a jacket they already owned, or trousers in a silhouette they had been circling around.
That was not a surprise in retrospect. It was exactly the value that professional styling delivers and that no discovery feed had ever attempted to replicate. We had built it into the product but had not foregrounded it in our own description of what Atorie does. The pilot pushed us to rewrite how we talk about the product and to invest more in the outfit-build layer rather than treating it as secondary to label search.
Quiz completion also surprised us. We had expected significant drop-off somewhere in the 40-question quiz, particularly in the sections that ask about specific garment categories you currently own. The median completion was question 36. Several participants commented that the quiz felt specific enough to hold their attention: they were not being asked vague questions about "style aesthetic" but concrete questions about actual garments and real occasions. The specificity maintained engagement even at length.
Three things we changed based on the pilot
The first change was adding what we call the "existing wardrobe" layer to the quiz. Before the pilot, the quiz was entirely forward-looking: where do you want to go with your style, what occasions do you need to dress for, what price range works for you. Participants told us repeatedly that what they already owned was shaping and constraining their choices in ways the original profile never captured. Someone with a wardrobe of structured dark pieces needed something different from label suggestions than someone building from scratch. We rebuilt the input layer to include questions about owned-garment categories, and that data now feeds directly into how outfit suggestions are assembled.
The second change was to the match score display. The original interface used percentage match scores as the primary visual element on each suggestion card. Participants found those numbers alienating. They did not know what a 91% match meant and did not trust it. What they wanted was to understand the attributes driving the match: why is this jacket showing up, what about my profile does it connect to. We moved the numeric score to a secondary position and made the attribute breakdown the primary explanation visible on the card. Usage data from the pilot's final two weeks, after we shipped this change to a subset of participants, showed higher engagement with the cards that showed attribute reasoning.
The third change was to the skip mechanic. Originally, skipping a suggestion had no visible consequence on the interface. The profile was updating internally, but the participant could not see it happening. Several participants described this as feeling like "shouting into a void." They wanted confirmation that their choices were being incorporated and a way to verify how their profile was shifting. We added explicit visual feedback when a skip is registered, and we built a profile summary view that shows your taste attributes and how they have moved based on your recent interactions. That view became one of the more-used features in the pilot's final week.
Where we were honest with ourselves about problems
Cold-start quality is the hardest open problem we have. The first session of matches, before any interaction data has updated the profile, is noticeably weaker than matches from session two onward. Several participants described their initial suggestions as "in the right neighborhood" but not accurate, and two of the professional stylists in the pilot used the word "generic." They were not wrong. Quiz answers give us a rich starting profile, but they cannot fully replace the signal that comes from a user interacting with actual suggestions. The gap between session-one quality and session-two quality is wider than we want it to be, and improving cold-start matching is the most active area of product work right now.
We also heard consistent feedback about the mobile experience. The outfit card format works well on desktop, where you can see multiple garments at once and understand how they relate. On a phone, the cards required more scrolling than participants wanted. This is a layout problem, not a content problem, and we are redesigning the mobile card format. We are not shipping anything until the pilot participants who flagged this issue have had a chance to test the revised version.
One thing we are not changing, despite some feedback to do so: the quiz length. Several participants suggested a shorter onboarding quiz. We understand the intuition. But the 40-question structure is earning the quality of match that generates the outcomes participants valued. We looked carefully at drop-off data and at outcomes by completion level. Participants who completed more of the quiz had substantially better match outcomes in week one than participants who stopped earlier. We are not saying a shorter quiz is wrong in principle. We are saying that for a discovery tool where first-session quality is already our biggest weakness, shortening onboarding would make that problem meaningfully worse. The right move is to improve cold-start quality from within, not to reduce the input that cold-start matching depends on.
What the pilot told us about who this product is for
Thirty-two of the 40 participants said at exit that they planned to continue using Atorie after the pilot. Of those 32, the strongest retention signal correlated not with fashion-industry background but with how frustrated participants had been with existing discovery tools before joining. The people who found Atorie most valuable were the ones who already knew that mainstream recommendation algorithms were not serving them and had a specific sense of what they wanted instead.
That is the population we are building for. Not everyone who buys clothes, but people who have already concluded that the default discovery systems are broken and who are actively looking for something that takes their taste seriously. Those people exist in larger numbers than we initially estimated. We are grateful that 40 of them spent four weeks testing a version of the product that still has real rough edges, and even more grateful that they told us where those edges were.