Advancing Maternal Healthcare: How LLMs are Bridging Information Gaps with Trust and Readability

Explore how leading LLMs like ChatGPT, Perplexity AI, and GeminiAI are evaluated for delivering accurate, readable, and culturally sensitive routine maternity advice, improving digital health literacy in underserved communities.

Advancing Maternal Healthcare: How LLMs are Bridging Information Gaps with Trust and Readability

The Imperative for Accessible Maternal Health Information

      Maternal health remains a critical global concern, particularly in regions where healthcare infrastructure is fragmented and socio-cultural factors hinder open medical consultation. Many expectant mothers, especially in diverse communities, often face hesitation in discussing sensitive pregnancy-related concerns with doctors due to stigma or privacy issues. This environment has led to a growing reliance on digital platforms, including Large Language Models (LLMs), as primary sources for everyday health queries. These AI models offer a potentially quick, empathetic, and privacy-preserving avenue for information, yet their reliability, readability, and cultural sensitivity in such a crucial domain demand rigorous evaluation.

      Recent academic research has underscored the potential of LLMs to transform access to health information, especially in contexts where traditional healthcare access is limited. This study, "Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice" (arXiv:2603.16872v1, 2026), systematically evaluates how effectively leading LLMs like ChatGPT-4o, Perplexity AI, and GeminiAI provide medically reliable, understandable, and culturally appropriate responses to common pregnancy questions, including prevalent myths and misconceptions. The findings highlight the significant role AI-driven tools can play in complementing existing healthcare services and empowering individuals with accurate information.

Evaluating AI for Trust and Clarity: A Multi-faceted Approach

      To assess the quality and reliability of LLM-generated responses against those from qualified medical practitioners, the research implemented a structured, multi-level evaluation framework. This methodology focused on key parameters crucial for practical health communication: readability, semantic similarity, and noun entity overlap. The study involved a comprehensive set of seventeen pregnancy-related questions, designed to cover topics from early gestation to postnatal care, specifically including local cultural myths.

      Readability was a primary focus, ensuring that information is accessible to individuals from diverse educational backgrounds. Linguistic metrics such as Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), SMOG Index, Dale–Chall Readability Score (DCRS), and Automated Readability Index (ARI) were employed. These metrics objectively assess text complexity, sentence length, word choice, and syllable count, translating them into scores indicating ease of comprehension. For example, a higher FRE score indicates easier reading, while a lower FKGL score suggests suitability for a broader audience.

      Beyond readability, the study delved into the content's meaning and contextual relevance. Semantic similarity was quantified using cosine similarity, a technique that compares the meaning of two text samples by analyzing their vector representations. A score close to 1 signifies high similarity in meaning between the LLM response and expert medical advice. To further measure the overlap in specific concepts and terminology, Jaccard similarity was used to assess noun entity recognition overlap, indicating how much shared terminology existed between the AI-generated and expert responses. This comprehensive approach allowed for a nuanced understanding of the LLMs' capabilities. ARSA Technology applies similar rigorous evaluation in developing custom AI solutions for mission-critical applications.

Key Findings: Pinpointing Strengths and Potential

      The detailed analysis revealed distinct strengths among the evaluated LLMs:

  • Readability: ChatGPT-4o emerged as the leader in generating the most readable and clinically coherent responses. It achieved the highest Flesch Reading Ease (FRE) score of 64.63, indicating content understandable to a middle school reading level (13-15 years old), which aligns well with public health communication standards. It also scored favorably in Flesch-Kincaid Grade Level (FKGL) at 7.53, further confirming its accessibility.
  • Semantic Alignment: Perplexity AI demonstrated the highest semantic similarity with expert medical responses, achieving a mean cosine similarity of 0.206. This suggests that while other LLMs might offer more readable text, Perplexity AI was better at conveying the core medical meaning in alignment with professional advice.
  • Contextual Overlap: ChatGPT also exhibited stronger contextual alignment, with Jaccard similarity scores ranging from 0.172 to 0.216, indicating a greater overlap in specific noun entities mentioned compared to other models. This suggests a better grasp of the specific concepts and terminology relevant to the queries.


      These findings collectively underscore the significant potential of LLMs, particularly ChatGPT and Perplexity AI, as valuable supportive tools. They can effectively address pregnancy-related myths, enhance maternal health communication, and improve digital health literacy, particularly in underserved communities where access to healthcare professionals may be limited.

Implications for Global Digital Health Literacy

      The growing engagement with LLMs for sensitive topics like pregnancy marks a pivotal shift in how individuals seek and receive health information. The ability of these models to provide quick, empathetic, and confidential responses offers a unique opportunity to overcome traditional barriers like social stigma and limited awareness. The study's focus on readability ensures that the information is not just accurate but also truly usable by a broad audience, including those with diverse educational backgrounds. This is crucial for fostering informed decision-making and promoting better health outcomes.

      The presence of widespread misinformation and culturally ingrained myths related to pregnancy, particularly in contexts with expanding internet usage, makes AI-driven tools even more vital. By providing reliable and culturally sensitive information, LLMs can play a critical role in combating the spread of inaccuracies and bridging information gaps. This aligns with ARSA Technology's vision of building the future with AI and IoT, deploying practical AI that delivers measurable impact. Our Self-Check Health Kiosk, for example, exemplifies how AI can make health screening more accessible and consistent, even in public facilities.

ARSA's Commitment to Responsible AI in Healthcare

      For AI to be truly beneficial in sensitive sectors like healthcare, issues of trust, safety, and accuracy cannot be compromised. ARSA Technology, with expertise in Artificial Intelligence and Internet of Things solutions, is dedicated to deploying production-ready systems that uphold these values. Our approach prioritizes full-stack AI engineering, ensuring that solutions are not only technically robust but also engineered for privacy, scalability, and operational reliability from concept to deployment. Our experienced team since 2018 consistently focuses on delivering solutions that translate complex technological capabilities into practical, real-world outcomes across various industries, including healthcare.

      Whether it’s through AI video analytics for safety and compliance or specialized AI solutions for health monitoring, ARSA understands that success hinges on precision and a deep understanding of operational realities. The rigorous evaluation of LLMs for maternity advice reinforces the necessity of such an approach, demonstrating that AI in healthcare must be deployed with careful consideration for user comprehension, cultural context, and medical accuracy.

      The findings from this research provide compelling evidence that LLMs have the potential to become invaluable adjuncts in global maternal health efforts. By systematically assessing their capabilities, we can better harness AI to support expectant mothers with reliable, accessible, and sensitive information, thereby enhancing digital health literacy and contributing to healthier communities worldwide.

      To explore how ARSA Technology can assist your organization in leveraging AI and IoT for enhanced operational efficiency and impactful outcomes, we invite you to contact ARSA for a free consultation.

      Source: Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice. Sai Divya Vissamsetty et al. arXiv:2603.16872v1 [cs.CL] 12 Jan 2026.