Executive summary
Cosmetic surgery buyers are unusual. Almost no one walks in cold. The average buyer researches for months, reads before-and-after galleries on their phone at midnight, watches consultation reels, and asks somebody they trust — sometimes a friend, increasingly an AI assistant. The decision is durable, expensive, sometimes concealed from a partner, and always irreversible. Whichever surgeon's name shows up in the recommendation shortlist tends to stay in the shortlist. Whoever isn't on it usually doesn't get considered at all.
This report measures who AI assistants name when NYC buyers ask about cosmetic surgery. The current published run contains 0 buyer-style questions across four AI assistants, sampled across 0 conversation-state buckets. It is split into a 0-response ranking layer (questions expecting a surgeon recommendation) and a 0-response editorial layer (questions where the appropriate answer is advice, triage, or a referral to the emergency room). Denominators come from one atomic, quality-checked scan run.
- The consensus tier is unusually narrow. Only 0 practices currently clear the confidence-interval threshold. Compared with adjacent verticals, cosmetic surgery in NYC is a two-name market at the top and a fragmented, sub-specialty-driven field below it.
- Per-assistant divergence is extreme. Some mid-tier surgeons are named almost exclusively by a single assistant — one practice in the current snapshot appears in Gemini responses fifty-five times and elsewhere twice. "AI-visible" is not one thing; it is four different things with four different denominators.
- Conversation state — not keyword intent — is the right axis. A buyer asking "who does a natural-looking deep-plane facelift" gets a different response shape than a buyer typing "I woke up with a hematoma," even though a keyword tool might file both under "facelift." Splitting the corpus along conversation state surfaces surgeon-citation behavior that a search-console-first methodology cannot see.
- Crisis and constraint questions route away from named surgeons. When the prompt is "I think my implant ruptured last night" or "my consult made me feel dismissed," AI assistants correctly refuse to substitute for medical judgment and route the buyer to emergency care, board directories, or second-consultation guidance. These are in our editorial layer and are explicitly excluded from the ranking; counting them would distort every surgeon's share.
The methodology underlying this report — versioned prompt sets, two-layer ranking/editorial split, honest confidence intervals, immutable snapshots, full verbatim source traceability — is published at viclaro.app/leaderboards/methodology. Every citation in the rankings below traces back to a specific AI response.
1. Why measure this at all
The visibility problem in cosmetic surgery has always been unusual. Buyers can't easily comparison-shop the way they can with divorce firms or wealth managers. Outcomes are hard to evaluate before the fact. Word-of-mouth still dominates, but the "word" increasingly comes from something that isn't a mouth.
Three behaviors have shifted since 2023. First, buyers who used to open Google to search "best rhinoplasty NYC" now open ChatGPT and describe what they actually want ("someone who does natural, ethnic rhinoplasty on Middle Eastern noses without changing character, ideally under fifty"). Second, buyers who receive a referral from a friend now paste that surgeon's name into an assistant to see whether the AI corroborates the recommendation — a validation step that didn't exist before. Third, buyers who have a bad first consultation now describe the consult in the assistant and ask who they should see for a second opinion. All three interactions rewrite the shortlist. None of them are visible in Google Search Console.
The consumer economics reinforce the shift. A rhinoplasty in NYC is $15,000–$40,000 out of pocket. A deep-plane facelift is $50,000–$150,000. Revision work costs more than primary work. Buyers who have already spent months on Instagram, RealSelf, and TikTok are asking the assistant to compress that research into a shortlist. Being on the assistant's shortlist is now a material distribution channel, and this report is the first reproducible measurement of who currently occupies it.
2. Methodology
The published snapshot uses the versioned prompt set " " with 0 buyer-style questions grouped into 0 conversation-state buckets. The buckets describe where the buyer is in the recommendation conversation, not what keyword they would type:
- Situation, Decision, Validation, Follow-up — the ranking layer: questions where the buyer is asking AI to evaluate or recommend surgeons.
- Advice, Crisis, Constraints, Learning — the editorial layer: emotional or medical states in which AI typically declines to name specific surgeons.
The two-layer split matters more here than in most verticals. Cosmetic surgery generates unusually loaded questions. "Am I too old for a facelift?" "My mom keeps saying I should do something about my eyes — should I?" "Is it vain to want this?" "My husband says don't do it — is he right?" These questions produce genuinely helpful AI responses, but the responses almost never name individual surgeons — and correctly so. Counting them would deflate every practice's share by pulling non-recommendation responses into the denominator. The two-layer split lets us measure both honestly: the ranking layer reflects who AI recommends when asked, and the editorial layer reflects what AI does instead when the buyer hasn't yet asked for names.
Every question was posed independently to Claude, ChatGPT, Gemini, and Perplexity, with multiple samples per (question × assistant) pair to capture how consistently each assistant returns the same names. There is no memory or conversation history carried in between samples, except for the Follow-up bucket, which explicitly stages a multi-turn exchange ("of those, which does en-bloc removal specifically?").
Each AI response was parsed into a JSON array of named practices. Names were normalized — honorifics stripped from de-duplication, plural clinic/practice suffixes collapsed, sub-locations mapped to the parent practice, board-certified suffixes preserved — and matched against a pool of NYC cosmetic-surgery entities. When an assistant mentioned a surgeon we hadn't indexed, we looked them up and added them so future scans would recognize the name.
Three statistics feed the ranking layer:
- Response mention rate — distinct eligible responses naming the practice ÷ ranking-layer responses. Rates are non-exclusive because one response can name multiple practices.
- Honest 95% confidence interval on the share. The interval widens with smaller N; the same share at a smaller sample size produces a wider band.
- Cross-assistant coverage — how many of the four assistants (out of 4) cited the practice at least once. A practice cited by all four is a structurally different signal from one cited only by Perplexity.
Tier classification: Consensus if the CI lower bound is clearly above the noise floor; Mid-tier if cited by at least two assistants with share above 3% but the CI still spans the floor; Long tail otherwise. Long-tail practices are still ranked, but the position should be read as ordering rather than as a statistical claim.
Snapshots are immutable. Every ranking is reproducible against the same prompt-set version and the same assistant families, within measured sampling variance. Every citation has a verbatim-AI-response permalink for research and licensing partners.
Bucket structure and per-version question counts are published at viclaro.app/leaderboards/methodology. Verbatim question text isn't published — the wording is part of what makes the Atlas measure conversation-state behavior rather than search-engine behavior.
3. The consensus tier — a two-name market at the top
Of the ten current top entries, only 0 clear the consensus-tier bar: their CI lower bounds are clearly above noise and they are cited across all four assistant families. Every other practice on the top-ten list has either meaningfully lower share or coverage from three assistants rather than four. That is a distinctive finding: cosmetic surgery in NYC is not a broad-consensus market. It is a two-name market at the very top, and then a long, fragmented tail.
What the consensus tier has in common. Both consensus-tier surgeons operate as named-partner practices with the surgeon's own name as the URL, not a corporate group name. Both publish extensive before-and-after galleries organized by procedure and demographic. Both have written or been quoted in mainstream press on named procedures — deep-plane facelift technique, ethnic rhinoplasty, revision work. Both have decades-long ownership of a specific procedure vocabulary that the assistants can quote back verbatim. That combination — named surgeon, procedure-specific gallery URL structure, quotable technical thesis, sustained press footprint — is what shows up in the top of every top-ten list we've built.
4. The mid-tier — sub-specialty niches and single-assistant dominance
Below the consensus tier, the picture fractures. Cosmetic surgery isn't a single service. A buyer asking about a deep-plane facelift is not asking the same question as a buyer asking about ethnic rhinoplasty, breast reduction, gender-affirming surgery, en-bloc explant, or gynecomastia. The mid-tier reflects that. Practices appear because they own a specialty lane — not because they're broadly recommended for cosmetic surgery generally.
Two patterns dominate the mid-tier:
- Single-assistant dominance. Several mid-tier practices are essentially one assistant's recommendation. One practice in the current snapshot is cited fifty-five times by Gemini and effectively nowhere else. Another is cited thirty times by Perplexity and essentially nowhere else. These are not "AI-visible" practices in a general sense. They are Gemini-visible or Perplexity-visible practices, and if a meaningful share of a firm's actual buyers use a different assistant, the practice does not exist in that buyer's shortlist at all.
- Sub-specialty positioning. Practices that publish focused content around a specific procedure — closed rhinoplasty, revision work only, transgender top surgery, en-bloc explant — surface almost exclusively on the questions that match their lane. Their share on the broad "top rhinoplasty surgeons in NYC" question is small; their share on "who does en-bloc breast implant explant with the shortest scar" is disproportionate. The cost of that positioning is breadth. The benefit is a defensible territory that the consensus-tier generalists don't compete for.
A separate observation: one of the ten current top entries is not a surgeon at all but a hospital system. Institutional recommendations behave differently from individual surgeon recommendations — they appear on questions where the buyer is hedging risk, asking about complication management, or has been referred by a primary-care physician. The presence of an institution in the top ten is not a mistake; it's a signal that a portion of buyer queries route to institutional risk-reduction rather than to individual expertise.
5. The editorial layer — where AI sends buyers in crisis, doubt, or hard constraints
The editorial layer is separated from the ranking because it does something the ranking layer can't show. When a buyer is in crisis ("I have a hematoma the night after surgery"), medical uncertainty ("does capsular contracture ten years later mean I need explant?"), emotional pressure ("my husband says don't"), or working under hard financial or discretion constraints ("I need this hidden from my partner" / "I have $6K total"), AI assistants typically do not list named surgeons. They route the buyer toward emergency care, second-consultation guidance, financing frameworks, insurance-covered procedures, or the American Board of Plastic Surgery's find-a-surgeon directory.
Three families of question sit in this layer:
- Post-operative crisis. Suspected infection, hematoma, seroma, capsular contracture, implant illness, wound dehiscence, difficulty breathing after rhinoplasty. The correct AI behavior on these questions is to refuse to substitute for medical judgment and route the buyer to the operating surgeon or emergency care. Every assistant tested does this correctly, and firm citations on these questions are rare and — where they do appear — usually cite the original surgeon by name because the buyer already named them in the question.
- Emotional and identity questions. "Am I too old?" "Is it vain?" "My mother thinks I should have something done." "My daughter wants a rhinoplasty at nineteen — should I let her?" AI assistants respond with framework, not names. This is arguably the highest-emotional-intensity slice of the buyer journey, and it is invisible in every conventional SEO methodology.
- Hard constraints. Budget capped below the market rate. Recovery hidden from a partner. Single-parent recovery windows. No general anesthesia tolerated. Medical tourism return-fixes. The AI response here is usually structural — how to think about the constraint — rather than a shortlist.
For practices, there are two responses to the editorial layer. The first is to publish content explicitly designed to be quoted on these questions: age-appropriate consultation content, informed-consent framing, complication triage decision trees, gender-affirming surgery ethics. Not to appear as the answer to a crisis question — that would be inappropriate — but to appear as a quotable authority when the assistant is describing the framework. The second is to recognize that the American Board of Plastic Surgery directory is now a distribution channel; a listing on it, with a complete profile, materially affects how AI routes early-funnel buyers to individual surgeons.
A ranking-only methodology would either drop these questions (losing the signal) or count them and dilute every practice's share. Splitting the layers lets us report both honestly.
6. Per-assistant divergence — why "AI visibility" isn't one thing
The four AI assistants tested do not behave the same way. Across the ranking-layer responses each produced for this category:
| Assistant | Eligible responses | Interpretation |
|---|
Beyond the denominators, the four assistants disagree qualitatively about who to recommend. Gemini concentrates its citations on a small number of practices with clearly named surgeons and procedure-specific URL structures. Claude spreads its citations more evenly and rewards practices that publish detailed technique explanations. ChatGPT — currently the highest-traffic assistant for cosmetic-surgery buyers — is the most conservative, often recommending institutions and board-certified directories alongside named surgeons. Perplexity, because it grounds every response in live web search, sometimes recommends practices with strong site presence but a thin training-data footprint; a practice with almost no citations in Claude or ChatGPT can be a Perplexity favorite.
The implication is that "are you AI-visible?" is not a single binary question. It has four different answers, each with its own denominator. A response with no ranked surgeon can reflect a genuine non-recommendation, a directory-only answer, an extraction miss, or a refusal — this report does not collapse those states into a single refusal statistic.
7. What separates the consensus tier from the long tail
The current snapshot contains 0 ranked entities: 0 consensus-tier, 0 mid-tier, and the remainder in the long tail. The following website and structural patterns are hypotheses for follow-up audits — not causal conclusions established by the recommendation panel alone:
- Surgeon-name domains. The consensus-tier practices live at URLs like drsamrizk.com or newyorkfacialplasticsurgery.com — the surgeon's name or a specific procedure phrase in the domain. Corporate group names, aggregator sites, and multi-surgeon collectives do not surface at the same rate. AI assistants pattern-match on "person + procedure + city," and a surgeon's own name in the URL is the strongest possible reinforcement.
- Procedure-specific landing pages, not service menus. Consensus-tier practices publish separate URLs for deep-plane facelift, closed rhinoplasty, ethnic rhinoplasty, upper blepharoplasty, deep-neck lift — each with its own long-form technique explanation and its own before-and-after gallery. Long-tail practices publish a single "Procedures" page that lists thirty options and links out. The assistants surface the specific page more consistently than the menu page.
- Before-and-after galleries with structured captions. Galleries that caption each photo pair with age range, ethnicity where relevant, procedure name, and time-since-surgery get quoted more often than galleries with just numeric labels. Assistants find the structured captions easier to lift into recommendation text.
- Named technique or thesis. "Deep-plane facelift with SMAS elevation," "preservation rhinoplasty," "en-bloc capsulectomy with intact removal" — surgeons who own a technique name in their marketing show up in queries about that technique. This is closer to trademarking a piece of vocabulary than to conventional SEO.
- Press footprint on procedure-specific stories. Not vanity press. Consensus-tier surgeons are quoted in Vogue, the New York Times style section, and specialty trade press on specific procedures. The article-plus-surgeon citation footprint carries into AI training data and reinforces the assistant's pattern-matching on the surgeon-procedure pair.
- Board certification, prominently displayed and schema-marked. ABPS or AAFPRS certification, displayed on landing pages and marked up with structured data, is over-represented in the consensus tier. It is not sufficient on its own — plenty of long-tail surgeons are board certified — but its absence is a near-universal disqualifier.
None of these signals is sufficient alone. The consensus-tier practices exhibit all six; the long-tail practices typically exhibit two or fewer.
8. Methodological honesty — what this report does not measure
A report of this kind must be explicit about its limits.
- We measure citation share, not surgical quality. A surgeon cited often by AI is not necessarily a better surgeon than one cited rarely. Complication rates, revision rates, board discipline records, and patient-reported outcomes are different questions. This report does not answer them.
- We measure four AI assistants, not all AI surfaces. ChatGPT in voice mode, Claude on mobile, Meta AI inside Instagram DMs, and various shopping-assistant wrappers may behave differently than the API-accessed assistants we test. The big four cover the dominant share of buyer behavior today, but the surface is fragmenting fast.
- Mid-tier rank movement is uncertain. The 95% confidence intervals on mid-tier practices span 2–3 percentage points. Reading rank changes from #7 to #5 between snapshots as a "rise" is not statistically supported at this sample size. Top-3 ranks are more stable; mid-tier requires larger N before movement is meaningful.
- We do not measure conversion. Being cited by ChatGPT does not necessarily produce a consultation. The recommendation channel exists. The conversion economics of that channel are still being measured.
- We do not adjudicate ethics. The Atlas does not measure whether a surgeon should be recommended, only whether they are. Buyers are responsible for validating credentials, reviewing before-and-after galleries in person, and asking the questions that the editorial layer's prompts are designed to model.
9. Conclusions
Three things are true at once:
AI recommendation visibility in NYC cosmetic surgery is now a measurable, ranked, reproducible distribution channel. The methodology behind this report — versioned questions, multi-sample scans, honest confidence intervals, immutable snapshots — produces ranks that survive re-running and can be defended to a sophisticated reader.
The channel is structurally narrow at the top and structurally winnable in specialty lanes. A two-name consensus tier is a durable equilibrium; displacing it requires the same combination of surgeon-name domain, procedure-specific pages, named technique, and press footprint that put those two names there. The mid-tier is a different opportunity: sub-specialty content targeted at a defined lane can produce durable single-assistant dominance without competing head-on with the consensus tier.
The conventional "AI optimization" pitch is partially wrong for this vertical. Most tools in the category measure a single assistant, a single prompt format, or a single "AI visibility score." Cosmetic surgery specifically defeats those simplifications. Per-assistant divergence is extreme. Editorial-layer questions dominate the emotional register of the buyer journey and don't produce firm citations at all. Sub-specialty questions surface entirely different names than general-category questions. A practice that optimizes for an averaged score risks optimizing for nothing in particular.
The full live index for this category — updated as we re-scan, with verbatim source traceability per citation — is at viclaro.app/leaderboards/nyc/medical/cosmetic-surgery. Methodology disclosure is at viclaro.app/leaderboards/methodology.