Request Proposal

When Medical Coding Becomes a Bottleneck: Lessons From a Real Clinical Study

From the Confidence Medical Affairs Desk by Ekaterina Shilova

Medical coding challenges are often discussed from a theoretical perspective, including inconsistent terminology, free-text entries, fragmented analyses, or delayed queries. However, the operational impact becomes much clearer when these issues unfold in a real study environment.

A recent clinical study demonstrated how quickly coding review can become a major operational bottleneck when data collection, EDC structure, and coding processes are not aligned from the beginning.

What initially appeared to be a manageable coding workload ultimately evolved into a highly resource-intensive process involving extensive manual review, thousands of queries, and significant downstream inefficiencies.

Study Overview

The study involved:

  • 2 Medical Monitors assigned to coding review at the end of the study 
  • 1,600 patients 
  • 90,000 coded items 
  • Average throughput of approximately 150 terms per hour 
  • 2,211 queries raised during coding review 

At first glance, these numbers suggested a heavy but operationally manageable workload. In reality, the challenges extended far beyond the volume of coded terms alone.

The primary issue was not simply the number of records requiring review. It was the quality and structure of the original data entering the system.

The Problems Started at Data Collection

The root cause of many downstream coding issues originated much earlier in the study lifecycle, specifically during data entry and EDC design.

The Indication field was completed manually by sites without:

  • predefined lists 
  • standardized terminology 
  • or links to Medical History, Adverse Events, or Concomitant Medications 

As a result, sites entered information using inconsistent terminology, vague descriptions, and non-standard medical wording.

At the same time, site personnel had not received sufficient guidance on what constituted medically meaningful and codable entries.

Site staff were not adequately trained to enter medically meaningful and codable information.

Although a coding convention had been established and actively used by the Medical Monitors, the effectiveness of those conventions was inherently limited by the quality of the incoming data.

This highlights an important operational reality in clinical research. Coding conventions alone cannot fully compensate for poorly structured source data.

How Free-Text Entries Created Downstream Risk

Once inconsistent or ambiguous data entered the EDC system, the downstream effects became increasingly difficult to control.

The study encountered:

  • frequent typos 
  • inconsistent terminology 
  • vague or medically unclear indications 
  • and widespread autocoding misclassification 

Examples included entries such as “antibacterial,” “blood pressure,” or “potassium” being used as indications.

Without proper context, autocoding systems interpreted these terms incorrectly.

→ “blood pressure” was classified as a procedure
→ “potassium” was interpreted as a laboratory parameter rather than a clinical condition

Although these may initially appear to be isolated coding issues, their impact extended much further. Misclassified indications fragmented analyses, increased manual review burden, and complicated interpretation across the study dataset.

In practice, the Medical Monitors spent substantial time manually reviewing and correcting entries that could have been prevented earlier through better EDC structure and site guidance.

The Operational Burden Escalated Quickly

As inconsistencies accumulated across the database, the operational workload increased significantly.

Despite having coding conventions in place, the team faced:

  • extensive manual recoding 
  • repeated clarification requests to sites 
  • prolonged back-and-forth communication 
  • and substantial review inefficiencies 

A total of 2,211 queries were ultimately raised to clarify site-entered data.

This created delays not only for coding review itself, but also for downstream data cleaning and final study deliverables.

The timing of the coding review further compounded the issue.

Why Late Coding Review Became So Costly

One of the most significant challenges was that coding review occurred near the end of the study.

By that stage:

  • the EDC structure was already finalized 
  • the database configuration could no longer be optimized 
  • and structural improvements were no longer feasible 

The EDC system could no longer be reconfigured, and improvements such as dropdown menus, field restrictions, or logical data linkages were not possible.

At the same time, site responsiveness was considerably lower late in the study lifecycle, making query resolution slower and more operationally burdensome.

This created a situation familiar to many study teams. By the time the coding issues became fully visible, the opportunity to efficiently prevent them had already passed.

Key Lessons From the Study

The study reinforced several important lessons regarding medical coding and operational planning in clinical trials.

First, coding quality begins with data collection, not with final coding review.

Free-text fields without guidance or structure introduce ambiguity that becomes increasingly difficult and resource-intensive to resolve later.

Second, coding conventions remain essential, but they are not sufficient on their own. Even the strongest conventions cannot fully compensate for inconsistent or medically unclear source data.

Third, EDC design plays a critical role in coding quality. Structured fields, controlled terminology, and logical links between datasets significantly reduce ambiguity and improve downstream consistency.

Finally, early Medical Monitor involvement is essential, particularly during study setup and database design phases.

Early Medical Monitor involvement is especially important during study setup and database design.

Medical Monitors help ensure that:

  • clinically meaningful terminology is captured appropriately 
  • structured data collection supports downstream analyses 
  • and coding approaches remain aligned with study intent and regulatory expectations 

The study also demonstrated that site training should never be considered optional.

Sites must understand:

  • what constitutes a valid indication 
  • how information should be entered 
  • and why medically meaningful terminology matters for downstream analyses and safety interpretation 

Conclusion

This study demonstrated a critical reality of clinical trial operations. High-quality medical coding cannot be retrofitted at study closeout.

By the time coding inconsistencies become visible during final review, the operational cost of correction is often already substantial, affecting timelines, resources, query burden, and overall data quality.

Effective coding strategy must begin at study startup through:

  • thoughtful EDC design 
  • structured terminology 
  • early Medical Monitor involvement 
  • aligned coding conventions 
  • and adequate site training 

Ultimately, high-quality coding is not simply a downstream data management activity. It is a foundational component of reliable clinical interpretation and overall study integrity.

Related News

Explore recent updates, announcements, and insights.

The Relationship between CROs and Pharmaceutical Companies

The following formula about clinical trials won’t appear in any textbook, but it is very real for small- to…

Current Trends in Clinical Research

Pharmaceutical research and development contributes significantly to today’s rising health care costs. Clinical research, a critical component of pharma…

Why These Are the Top Countries for Conducting Clinical Trials

Clinical trials are currently being held on every continent except Antarctica. However, emerging countries have long been rising to…