[비즈한국] While AI is transforming the drug development paradigm for biopharmaceutical companies, its role in the medical field is expanding beyond simply assisting doctors' diagnoses to generating clinical findings necessary for medical judgment. With the emergence of generative AI that does not stop at identifying abnormalities in medical images but even drafts initial interpretation reports, the so-called ‘AI Healthcare 2.0’ era is now in full swing.
In April, the medical AI company Deepnoid received regulatory approval from the Ministry of Food and Drug Safety for its generative AI-based chest X-ray solution, ‘M4CXR.’ On the 3rd, the company also signed its first supply contract with Pusan National University Hospital, accelerating its commercialization.
Pusan National University Hospital is a tertiary referral hospital with over 1,200 beds and a regional medical hub. Given that a large domestic hospital has introduced a generative AI-based image interpretation solution into actual clinical practice, it remains to be seen whether this will act as a catalyst for other medical institutions to follow suit.
As generative AI automates a significant portion of the interpretation process, clinical productivity is expected to improve substantially. However, as AI evolves beyond a simple diagnostic aid to generating findings and presenting information for clinical judgment, the question of who should be held responsible for AI errors is emerging as a critical issue.

Beyond Diagnostic Assistance to Writing Reports
Existing first-generation medical AI, such as that from Lunit and Vuno, assisted medical staff by highlighting areas suspected of having lesions on images and providing probability values. Medical staff would then confirm the AI-identified lesions and make a final judgment based on their own clinical experience and the patient's condition. In this sense, AI was merely a tool to aid the doctor's decision-making.
The M4CXR, based on generative AI, has gone a step further. It utilizes LMM (Large Multimodal Model) and CoT (Chain of Thought)-based reasoning to analyze chest X-ray images and generate initial interpretation reports in natural language. In other words, it does not just signal "there is an abnormal area"; it summarizes the features found in the images into sentence-based findings that medical staff can utilize.
Compared to the general-purpose Large Language Model ChatGPT, the accuracy of lesion localization has more than doubled while the hallucination rate has dropped below 0.2%. Consequently, the medical community currently assesses the generated output to be at the level of a second- or third-year resident. However, the accuracy and completeness could be further improved through continuous learning based on real-world clinical data.
The medical community expects such generative AI to significantly reduce the burden on radiologists. In particular, since about 95% of chest X-rays in general health checkups are currently deemed normal, an environment where AI filters these normal images with high accuracy would allow medical staff to focus more on images that potentially contain abnormalities.
Kim Sung-hyun, Director of the Human Radiology Center, evaluated M4CXR by stating, "The fact that generative AI can pre-filter 95% of normal cases with high precision reduces the overall workload for radiologists by that much. Since it even drafts the reports, doctors only need to provide final review and approval. We finally have an AI solution worth paying for."
The competitiveness of generative AI medical devices, therefore, depends not just on how well they identify lesions, but on how much they can actually reduce the workload of medical staff.
"Zero Labor but 100% Responsibility?" The Paradox of the Autonomous AI Era
If AI only identifies lesions in images, the doctor takes on the role of making the final judgment. However, if AI analyzes images and even generates interpretation findings, the doctor's role becomes focused on reviewing and approving the AI-generated results.
In the U.S. medical community, discussions regarding the responsibility issues stemming from these changes are in full swing. The American Medical Association (AMA) distinguishes medical AI into concepts like assistance, augmentation, and autonomy, emphasizing the principle that AI should be used under human supervision and responsibility rather than replacing the judgment of medical staff.
However, the problem is that as the AI's role grows, the boundaries of responsibility become increasingly blurred. Critics point out the contradiction that the moment one describes delegating work to AI, it ceases to be a simple tool and becomes an acting agent, yet the responsibility is still placed on the doctor.
Choi Yoon-seop, CEO of Digital Healthcare Partners (DHP), noted, "Simple medical devices like stethoscopes or thermometers are not objects to which a doctor delegates work. The concept of delegation only holds for entities with the capacity for action, like residents or interns, and some have expressed that this now applies to AI as well."
If AI has evolved to the level of performing the doctor's tasks, who should be held responsible for the results? Expecting a doctor to be able to predict and control the AI's entire decision-making process and potential errors when a medical accident occurs is an entirely different matter.
This shift is also appearing in U.S. medical AI regulations. While the reimbursement system operates on the premise that AI performs a significant portion of the medical act, the discussion regarding legal responsibility remains centered on the role and accountability of human medical professionals.
The U.S. Centers for Medicare & Medicaid Services (CMS) set the physician work value at '0' when calculating the reimbursement for 'CPT 92229,' a medical code for retinal disease detection using autonomous AI. This implies that the time, skill, effort, and clinical judgment required by a doctor to perform this act are not separately calculated as physician work.
On the other hand, the legal responsibility for medical accidents remains complex. Even if AI is deeply involved in medical acts, the current U.S. medical legal liability framework is fundamentally centered on the judgment and actions of licensed medical professionals. Consequently, there is growing debate over who—the medical professional, the institution, or the AI developer—should bear the responsibility and to what extent when AI makes an incorrect judgment.
Ultimately, even in the U.S., while the domain where AI participates in medical practice is expanding, the standards for compensation and responsibility are changing at different speeds.
This is why critics argue that if a doctor is not compensated for their direct labor because AI handles most of the work, yet still has to bear the burden of responsibility for AI errors, there is a clear imbalance between accountability and reward.
Kim Sung-hyun, Director of the Human Radiology Center, also pointed out, "As tasks performed independently by AI increase, if doctors are held responsible when a medical accident occurs, there will be backlash, with people asking, 'Why should the doctor be responsible when they didn't even perform the interpretation themselves?' There is currently no national system or social consensus on this at all."
Kwon Min-ji, a partner attorney at the law firm Do-A, stated, "Even if AI generated the findings, if a misdiagnosis occurs after a doctor's final approval, the current law holds the doctor who made the final decision responsible, not the assistant. It would be judged based on whether the doctor breached their duty of care, considering the medical standards, clinical environment, and the specific nature of the medical act at the time."