Sitting with someone through a hard moment is one of the core skills in clinical mental health and social work, the ability to stay present with a person's pain without rushing to fix it or looking away from it. Whether that skill can actually be learned, practiced, and delivered through a video screen is a real question facing training programs, not a settled one. It matters because so much clinical training and ongoing practice has moved to remote and hybrid formats, and the answer shapes how an entire generation of clinicians will be prepared to do this work.
Most people's first instinct is that presence needs a room, that something about sharing physical space allows a level of attunement a screen can't reproduce. That instinct runs into a body of outcome research that is more encouraging than most people expect, and the gap between the intuition and the evidence is the part worth sitting with directly.
The Intuition That Presence Needs a Room
Video conversation asks the brain to do work that in-person conversation doesn't require, constantly monitoring a small frame for cues that would register automatically in a shared space. Research on the strain of video interaction has traced this to what researchers call nonverbal overload, a combination of sustained close-up eye contact, the effort of sending and reading cues without the context of a full room, and the constant self-monitoring that comes from seeing your own face on screen throughout the conversation. If reading another person accurately already takes more effort over video, it seems reasonable to wonder whether something as subtle as therapeutic attunement survives the trip through a webcam at all, and plenty of experienced clinicians assumed for years that it wouldn't, treating the physical room as part of the treatment itself rather than just the setting for it.
Training programs have generally responded to that uncertainty by hedging rather than picking a side, building curricula where some coursework happens remotely and certain skills are still practiced in person with a supervisor watching directly. This is precisely what hybrid MSW programs name openly, treating the question of which skills need a shared room as a judgment made explicitly rather than left to convenience or cost, with specific relational skills, like reading a room during a family session or managing a moment of acute distress, still built through in-person practicum hours even as coursework and supervision move online.
What the Outcome Research Actually Shows
The caution built into hybrid training makes the outcome data more surprising, not less. Multiple controlled comparisons have found video-delivered therapy produces clinical outcomes and levels of therapeutic alliance that are statistically indistinguishable from in-person sessions, a pattern that has held up across depression, anxiety, and trauma-focused treatment alike, and across different age groups and presenting conditions.
Leslie Morland, who directs a regional telemental health program at the VA, has pointed out that clinicians consistently perform as well over video as they do face to face, and that the belief in a required physical presence persists more from professional habit than from what the data actually shows. That's a genuinely uncomfortable finding for anyone who assumed the room itself was doing essential work, since it suggests the skill travels further than the intuition predicted, and further than most training programs have been willing to fully act on.
What This Tension Actually Means
Holding both facts at once, that video conversation taxes the brain in ways in-person conversation doesn't, and that measured outcomes barely move as a result, suggests that whatever makes sitting with someone effective doesn't depend on the specific mechanism most people assume it does. The skill may have less to do with a shared room and more to do with a clinician's capacity to stay attentive and grounded regardless of the medium carrying the conversation, a capacity that can apparently be built and practiced through a screen more reliably than expected, and one that a supervisor can still observe and correct just as closely over video as they could standing in the same room.
Training programs that keep a foot in both formats aren't necessarily hedging out of doubt about the evidence itself. They may simply be acknowledging that a skill this personal is worth practicing in more than one setting before anyone fully trusts it either way, and that caution and confidence can reasonably coexist while the field keeps watching what the data continues to show, revising how much weight each format deserves as more evidence accumulates rather than settling the question once and moving on.

