Skip to content
82
Categories
Keynote incl. Free Communication

Research Technology and AI

- , Deck 3-4

Schedule Slot

Research Technology and AI

82
Categories
Keynote incl. Free Communication

Research Technology and AI

- , Deck 3-4
  1. Clinical application of AI in relation to head and neck reconstructive surgery

    Presentation time:
    20 min

    Speaker: Joep Kraeima

  2. AI-Enabled 3D Planning and Projection-Guided Breast Reconstruction

    Presentation time:
    20 min

    Speaker: Stefan Hummelink

  3. Exploring the Role of Artificial Intelligence in Decision-Making for Breast Reconstruction after Malignant Breast Pathologies

    Presentation time:
    6 min
    Discussion time:
    2 min

    Abstract Presenter: L. De Pellegrin

    Objective

    Breast malignancies pose complex diagnostic and reconstructive challenges, particularly when individualized strategies are required. Optimal management relies on multidisciplinary expertise and clinical experience. Artificial intelligence (AI) has emerged as promising tool to support clinical workflows by integrating patient-specific data to enhance diagnostic accuracy and therapeutic planning. This study evaluates the potential role of AI in supporting reconstructive decision-making in malignant breast pathologies.

    Methods

    Six clinical scenarios were created including standardized clinical information (medical history, clinical examination findings, and photo documentation) along with predefined reconstructive treatment options. The cases were submitted to ChatGPT-5.2 and also distributed to board-certified plastic surgeons and residents. Both, AI model and survey participants were instructed to choose the most appropriate therapeutic strategy.

    Results

    In Scenario 1 (extensive DCIS, low BMI), AI recommended subpectoral direct-to-implant reconstruction, whereas 50% of participants preferred autologous reconstruction and 25% selected two-stage implant reconstruction. In Scenario 2 (BRCA mutation), AI proposed bilateral autologous reconstruction, while participants showed heterogeneous preferences, most commonly bilateral mastectomy with two-stage implant reconstruction (38%). In more advanced or irradiated cases (Scenarios 3–6), agreement increased. In Scenario 3 (extensive tumor volume) , 63% of participants concurred with AI’s recommendation. In Scenario 4 (secondary reconstruction) 43% selected autologous reconstruction. In Scenario 5 (prior BCS+RTX), 43% agreed with AI, while 29% opted for implant-based or using fat grafting. In Scenario 6 (BCS+CTX+RTX), concordance was highest, with 63% favoring autologous reconstruction.
    Overall, concordance between AI and participants was greater in complex oncologic settings and lower in prophylactic or early-stage scenarios, where human decision-making appeared more heterogeneous.

    Conclusion

    AI model provided clinically plausible recommendations with higher concordance with expert opinion in complex cases. Greater variability was observed in early-stage or prophylactic settings. While not yet suitable for independent decision-making, further refinement, training, and expert validation could enhance AI’s role in future clinical applications.

  4. Simulated Microgravity Enhances Post-Exposure Proliferation in Human Adipose Stem Cells for Tissue Engineering Applications

    Presentation time:
    3 min

    Abstract Presenter: G. Z. Zinner

    Objective

    Microgravity profoundly alters cell behaviour, offering a unique microenvironment for mechanobiological studies and next-generation regenerative medicine applications. Random positioning machines (RPM), represent an established ground-based tool to simulate microgravity conditions. Human adipose-derived stem cells (hASCs) are of particular interest for regenerative medicine given their accessibility, multipotency, and paracrine potential. This study aimed to investigate the short-term response of primary hASCs to 72-hour RPM culture, focusing on survival, differentiation status, three-dimensional self-organization, and post-exposure proliferative recovery.

    Methods

    Primary hASCs isolated from human adipose tissue donors were cultured under standard conditions (37°C, 5% CO₂) or subjected to 72 hours of simulated microgravity using RPM bioreactor. Cell viability was assessed by live/dead fluorescence assay and metabolic activity by alamarBlue assay. Differentiation status was evaluated by flow cytometry using a panel of mesenchymal surface markers. Actin cytoskeleton organization and nuclear morphology were assessed by immunofluorescence with phalloidin staining. Proliferative activity following return to standard gravity was quantified immediately and at 48 hours post-exposure by Ki67 immunofluorescence staining.

    Results

    RPM-cultured hASCs maintained viability comparable to static controls at 72 hours, as confirmed by both live/dead and AlamarBlue assays. FACS analysis revealed no significant change in stem cell markers expression, indicating no phenotype change under these conditions. Phalloidin staining demonstrated pronounced actin cytoskeletal reorganization, characterized by cortical actin redistribution, and reduction in nuclear size compared to controls. Ki67 immunofluorescence revealed significantly enhanced proliferative activity in RPM-conditioned hASCs at 48 hours after return to standard gravity.

    Conclusion

    Short-term simulated microgravity via RPM induces substantial mechanoadaptive responses in primary hASCs, including spontaneous 3D aggregate formation, actin cytoskeletal reorganization, and accelerated post-exposure proliferation, without significant cytotoxicity or differentiation commitment. These findings suggest that microgravity conditioning may represent a promising strategy to modulate hASC behavior for tissue engineering and regenerative medicine applications.

  5. Artificial Intelligence as a Surgical Advisor before a DIEP Breast Reconstruction. A Blinded Comparative Study of Three Large Language Models

    Presentation time:
    6 min

    Abstract Presenter: S. Holm

    Objective

    Large language models (LLMs) are increasingly used in clinical communication, but their accuracy and readability in patient education remain unclear. This study compared three LLMs for preoperative counselling before a DIEP breast reconstruction.

    Methods

    A total of 40 frequently asked preoperative questions regarding DIEP breast reconstruction were collected and categorized using the BREAST-Q framework. These were submitted in English to three LLMs: ChatGPT, Gemini and Copilot (anonymised as Model A-C). Each question was submitted to all three models and the responses were anonymized. An expert panel of eight board-certified plastic surgeons from both Europe and USA. Ratings were made of a 5-point Likert scale for accuracy, informativeness and readability. Together with a general evaluation (easiness, problematic content, incorrectness) and information-material specific evaluation (relevance and lowest reading level).

    Results

    Significant differences were found between models across all domains. ChatGPT achieved the highest accuracy (p=0.019), Copilot was the most informative (p = 0.041), and both ChatGPT and Copilot produced more readable responses than Gemini (p < 0.001). Copilot had fewer problematic statements, while Gemini generated text at the simpleast reading level but with lower accuracy. Agreement among raters was strong for accuracy (κ = 0.96) but weak for qualitative domains.

    Conclusion

    Each LLM showed distinct strength ChatGPT produced the most accurate answers, Copilot the most informative, and Gemini the simplest language. No model was uniformly superior. These findings support supervised, task-specific use of LLMs in patient education for breast reconstruction.

  6. AI-assisted ambient dictation reduces documentation time in a multilingual Swiss hospital setting: a proof-of-concept study

    Presentation time:
    6 min
    Discussion time:
    2 min

    Abstract Presenter: M. Gladysz

    Objective

    To compare the efficiency and documentation quality of four clinical documentation workflows -- including AI-assisted and traditional methods -- used by a native and a non-native German-speaking physician in a multilingual Swiss tertiary hospital setting.

    Methods

    In this proof-of-concept study (Plastic and Hand Surgery, Cantonal Hospital Aarau), two physicians documented 12 simulated patient encounters across four workflows: (1) traditional dictation to secretary; (2) Dragon Medical One speech recognition; (3) Whisper v3 transcription with GPT-based note generation; (4) ambient dictation recording full appointments with AI transcription and GPT processing. Simulated patients included standard German, Swiss German dialect, and non-native German speakers. Physician time was measured by an independent observer. Note quality was scored using a modified PDQI-9 (scale 0-50) by three large language models.

    Results

    Workflow 4 (ambient dictation) produced the shortest documentation times: median 99s (native) and 84s (non-native), versus 294s and 586s for workflow 2 (speech recognition). Non-parametric ANOVA showed significant workflow-language interaction (P=.001) and workflow main effect (P<.001). Workflow 4 was significantly faster than workflow 2 for both physicians (adjusted P<.001). For the non-native speaker, workflow 4 versus workflow 1 did not reach significance (P=.077). All workflows achieved high quality (mean PDQI-9 >46/50) with clinically negligible differences (<1 point). Inter-rater reliability among AI evaluators was poor (Krippendorff alpha=-0.433).

    Conclusion

    AI-assisted ambient dictation demonstrated significant time savings -- up to 8.4 minutes per note for the non-native speaker -- without compromising documentation quality in a multilingual Swiss setting. These findings suggest AI scribes can help alleviate documentation burden and improve equity for non-native speaking physicians. However, AI-based quality scoring showed systematic disagreement, highlighting the need for human evaluators. These preliminary results support further validation with larger samples, real patient encounters, and secure data-processing frameworks before routine clinical deployment.

  7. Sensitivity of Photoplethysmography (PPG) to Early Ischemic and Congestive Changes in Free Flap: An In Vitro Phantom Study

    Presentation time:
    6 min
    Discussion time:
    2 min

    Abstract Presenter: H. Kodama

    Objective

    Free flap reconstruction carries a critical risk of vascular thrombosis, making early detection vital for successful salvage. Since clinical assessment remains subjective and labour-intensive, objective continuous monitoring is essential. Photoplethysmography (PPG) offers a cost-effective, non-invasive alternative, but evidence regarding its diagnostic performance via pulse wave morphological evaluation in free flaps remains scarce. This study aims to determine measurement depth limits and validate PPG’s diagnostic capacity for early vascular compromise using a customisable in vitro free flap model.

    Methods

    Custom silicone phantoms, accurately mimicking the mechanical properties of human flap pedicles, were embedded in a tissue-mimicking matrix at varying depths from 3 to 21 mm. Intravascular pressures were precisely controlled within an in vitro cardiovascular circuit to simulate normal, early ischemic, and early congested states. Reflectance PPG signals (Red and Infrared) were acquired and preprocessed. A comprehensive set of 47 morphological features was extracted from the pulse waves, and their diagnostic performance to differentiate hemodynamic states was evaluated using Receiver Operating Characteristic (ROC) analysis and Area Under the Curve (AUC).

    Results

    Signal quality assessment restricted reliable morphological evaluation to depths up to 15 mm. ROC analysis demonstrated that Intensity and Area-based parameters were consistently the primary diagnostic indicators across all viable depths, displaying a clear trend of Normal ≥ Congestion > Ischemia. Furthermore, depth-specific feature performance was identified: Time-based features were effective secondary markers at 9 mm, whereas derivative parameters (Slope and Second Derivative PPG) emerged as highly sensitive key indicators at 15 mm due to amplitude attenuation. Red light features were effective only at a superficial depth of 3 mm.

    Conclusion

    A physiologically calibrated in vitro phantom capable of replicating ischemic and congested states was successfully developed. Through comprehensive morphological analysis, PPG’s diagnostic capacity to accurately differentiate arterial failure from venous obstruction was verified. Identifying specific pulse wave feature variations across different depths establishes a robust framework for developing objective, continuous PPG monitoring systems for free flap assessment.

  8. Application of machine learning in facial palsy diagnosis, prognosis and dynamic functional assessment: A systematic review and meta-analysis

    Presentation time:
    6 min
    Discussion time:
    2 min

    Abstract Presenter: K. Kiew

    Objective

    Facial paralysis (FP) causes functional and psychosocial morbidity. Grading systems such as House Brackmann and Sunnybrook are subjective with variable inter-rater reliability. Machine learning (ML) may enable FP detection, severity grading, and recovery prediction using objective, quantitative assessment through image, video, and biosignal analysis.

    Methods

    A PRISMA-guided systematic review and meta-analysis were performed (PROSPERO CRD42024556789). Searches of PubMed, Embase, Scopus, and IEEE Xplore to December 2025 identified studies using ML for FP detection, grading, or recovery prediction. Data were extracted on study design, modality, model architecture, validation, and quantitative performance metrics (F1, AUC, accuracy). Risk of bias and reporting were assessed using PROBAST AI and TRIPOD AI, and evidence certainty with GRADE AI.

    Results

    Forty six studies (347876 image/video frame samples) were included. Pooled diagnostic and grading performance was high (F1=0.93 [0.90-0.96], AUC=0.94 [0.91-0.96]) with high sensitivity (0.91) and specificity (0.90). Transfer learning and hybrid deep architectures achieved the best overall accuracy (F1=0.97 [0.94-0.99], AUC=0.98 [0.95-0.99]), outperforming primary deep learning and ML models (p=0.011). Vision-based models involving image and video exceeded EMG biosignal approaches in accuracy (F1=0.94 vs 0.78, p<0.0001). However, static image models outperformed video models (F1=0.94 vs 0.91, p=0.003). Correlation between ML-derived and clinician severity scores was strong (r=0.80 [0.70-0.87], ICC=0.89 [0.85-0.93]). ML models grading FP severity and detecting FP achieved F1 of 0.93 (0.90-0.96) and 0.92 (0.86-0.96). Outcome prediction performance was lower with F1 of 0.80 (0.70-0.89) and AUC of 0.84 (0.76-0.90). Synkinesis and symptom screening demonstrated high accuracy (F1 0.98 [0.96-0.99]), but evidence was limited. Heterogeneity was high and external validation was rare, reducing certainty of evidence.

    Conclusion

    ML demonstrates high apparent accuracy and strong agreement with clinicians for automated FP assessment, particularly with transfer learning and vision-based models. However, further external validation and evaluation on patients are required. ML-based facial grading could standardise evaluation, guide intraoperative decision-making and rehabilitation monitoring, and move facial reanimation toward an optimised, data-driven approach.

  9. Discussion

    Presentation time:
    15 min