HomeResources › Webinar Questions Answered

The Questions We Did Not Reach, Answered in Writing

Eleven questions from the live session on clinical evaluation under EU MDR, 18 August 2026

The hour ran out before the chat did. These are the questions that were written during the session and never reached, answered here in full. Every question has been stripped of the name, the employer and the device that came with it, so what remains is the regulatory problem itself.

Watch the Full Session →

Who is answering, and who is not. The answers below are written by Dr. Hatem Rabeh, clinical evaluation expert and founder of Clinical Evaluation Navigator. They are not the guest reviewer’s answers, they are not a notified body position, and they are not an approval, a clearance or a decision on any device. Where the Regulation or a guidance document says something directly, it is cited so you can read it yourself.

What is deliberately absent. No names, no employers, no device names, no file under assessment. Several of these questions arrived with a described device attached. Those details have been removed and the question generalised, because the useful part is the principle and the identifying part helps nobody.

Acceptance criteria when the state of the art is thin

Three questions circled the same problem: the literature does not give you a number, and the criteria still have to be written before the data is analysed.

1. Safety and performance are measured in different units, so how can they be compared?

Asked live: safety usually appears as an occurrence rate, performance as a quantitative difference on a continuous scale.

Short answer: they are not compared with each other. Each is compared with the state of the art for that same endpoint, in that same unit.

The comparison a clinical evaluation makes is never safety against performance. It is your device against the state of the art, endpoint by endpoint. A complication rate is compared with the complication rate reported for the same procedure in the same population. A functional score is compared with the same score in the same population. The units differ because the endpoints differ, and that is not a problem to solve.

What does have to be solved is the shape of each criterion. A safety criterion needs a rate with a denominator that is stated, not implied: events per patient, per implant, or per patient-year, and which one it is. A performance criterion needs a central value and a range from the same pool of studies, and a stated direction, meaning which way is worse. Where a change from baseline is used, define the sign explicitly, because a lone negative number in a results table with no definition is a deficiency on its own.

One rule holds both together: the criteria are set in the clinical evaluation plan, from the state of the art, before the data is analysed. MDR Annex XIV Part A, section 1(a) requires the plan to contain “an indicative list and specification of parameters to be used to determine, based on the state of the art in medicine, the acceptability of the benefit-risk ratio for the various indications and for the intended purpose or purposes of the device”. If they are written afterwards, no amount of unit consistency will save them.

2. A novel device with no reliable benchmark: can the investigation endpoints become the acceptance criteria?

Asked live: the literature is not consistent enough to set a quantitative threshold, so the temptation is to reuse the clinical investigation endpoints, which were never designed as thresholds.

Short answer: no. That is circular, and it is visible from the outside. There are four honest ways out.

The order the Regulation implies runs one way only. The state of the art establishes what outcome level is achievable with existing treatment; those levels are selected as the acceptance criteria in the clinical evaluation plan; the investigation then measures against them. Deriving the threshold from the endpoints of your own study reverses that order, and it tells a reviewer the threshold was fitted to what you chose to collect. It is the same defect as fitting the threshold to the result, one step earlier.

When the literature genuinely will not give a number:

  • Widen the state of the art before concluding it is thin. It is not only similar devices. It is the alternative therapies for the same clinical condition, including drug treatment, surgery, and conservative management. Very often the benchmark exists, in the clinical literature of the condition rather than the device literature.
  • Use a qualitative or directional criterion and label it as one. Non-inferiority in direction, or a stated minimum clinically meaningful change taken from a cited source, is defensible. An invented number is not.
  • Declare the gap and give it to PMCF with a date and an owner. A named data gap reads as control. An undeclared one reads as an oversight.
  • Keep a priori assumptions in their own box. A value used to power a study is an assumption for the sample-size calculation, and it belongs in the plan labelled as exactly that. It is not an acceptance criterion, and the two must not appear in the same column.

On how this gets reviewed: the justification is read for whether the source of each threshold is traceable and whether the manufacturer noticed the weakness first. A file that names its own thin evidence and routes it to a real post-market commitment is in a far better position than one presenting a confident number with no origin.

3. One benchmark for the device, or one per indication?

Asked live, in the context of a device used across several indications where the existing benchmarks were established without stratification.

Short answer: stratify wherever the clinical outcome genuinely differs by indication, and say plainly what you did where the source data would not let you.

The test is clinical, not statistical: would a clinician expect a different outcome level in this population? If the answer is yes, a single pooled benchmark hides the worst subgroup, and the pooled result can pass while one indication fails. The intended purpose defines the scope of the evaluation, so if the intended purpose covers several indications, the evidence has to hold for each of them.

In practice this usually lands on a two-level structure. One headline endpoint for the device as a whole, and stratified acceptance criteria wherever the state of the art itself reports by indication. Where the published pool does not stratify, you cannot invent the strata: use the pooled benchmark, state in the plan that the source data is unstratified, and record it as a limitation rather than leaving the reader to discover it.

Where a substantial amount of your own data has accumulated, the useful move is to analyse by indication even when the criterion is pooled. If one subgroup is drifting, you want to be the one who found it.

Equivalence

Three questions, and all three came down to the same thing: what a difference costs you.

4. What is equivalence, and is it mandatory?

Asked live, as a request to go through it again from the beginning.

Short answer: it is not mandatory. It is one route to clinical data, and you only need it if you intend to rely on another device’s clinical data.

Equivalence means demonstrating that your device is equivalent to a device whose clinical data already exists, across three dimensions that must all hold at once: technical, biological and clinical (MDR Annex XIV Part A, section 3, elaborated in MDCG 2020-5). Failing any one of the three ends the claim; there is no partial credit and no averaging across the three.

Equivalence is not similarity, and the distinction decides what the data can be used for. A similar device informs the state of the art and gives you your benchmark values. It does not give you clinical data for your device. An equivalent device gives you clinical data you may use as if it were your own, which is why the bar is set where it is.

Two conditions catch people out. For implantable and class III devices, where the equivalent device is not your own, Article 61(5) requires a contract giving you full ongoing access to the technical documentation of that device. And the equivalent device has to be CE marked. A claim built on a device you cannot fully document is not a claim.

5. Can an equivalent device from the same manufacturer, used at a different anatomical site, carry the evidence?

Asked live: a class IIb implantable with no device-specific clinical data, considering an equivalent device from the same manufacturer placed in a different anatomical location, alongside similar-device data.

Short answer: a different anatomical site is normally a break in clinical equivalence, and same-manufacturer does not repair it.

The wording is in the Regulation itself. MDR Annex XIV Part A, section 3 defines the clinical characteristics as: the device “is used for the same clinical condition or purpose, including similar severity and stage of disease, at the same site in the body, in a similar population, including as regards age, anatomy and physiology; has the same kind of user; has similar relevant critical performance in view of the expected clinical effect for a specific intended purpose”. A different anatomical location goes directly at the site criterion, and the burden then sits with you to show that the difference produces no clinically significant difference in safety and performance. That has to be shown with data. An argument that the mechanism is the same is not data.

Being the same manufacturer helps with one thing only: it removes the Article 61(5) contract problem, because the technical documentation is already yours. It does nothing for the equivalence itself. The two are often conflated in submissions and they are entirely separate obstacles.

Practically, a class IIb implantable with no own clinical data, relying on an equivalent device at a different site, should expect that claim to be challenged. The realistic paths are a device-specific PMCF with a genuine data-collection commitment written into the plan, or a clinical investigation. Deciding that early costs far less than defending it late.

6. How much technical design difference is accepted in an equivalence claim?

Asked live, for a class IIb device with limited literature where equivalence has to carry most of the evidence.

Short answer: there is no percentage, and asking for one is the wrong question. Every difference is assessed for its clinical effect, one at a time.

No source sets a tolerance, and none gives a percentage. The standard is in MDR Annex XIV Part A, section 3: the characteristics “shall be similar to the extent that there would be no clinically significant difference in the safety and clinical performance of the device”. MEDDEV 2.7/1 Rev. 4, Appendix A1 puts the same test on each difference: they “need to be identified, fully disclosed, and evaluated; explanations should be given why the differences are not expected to significantly affect the clinical performance and clinical safety of the device under evaluation”. So the question is never how different the two devices are. It is whether this particular difference changes what happens to the patient.

The differences that most often break a claim are the ones that change a physical interaction: load transfer and mechanical behaviour, the material in contact with tissue and its degradation, geometry that changes fixation or fit, anything altering dose or energy delivered, and changes to how the user operates the device.

What survives review is a difference-by-difference table where each row states four things: the difference itself, the potential clinical effect, the specific evidence that closes it, and any residual uncertainty routed to PMCF. What does not survive is a table with a column of “no clinical impact” and no source behind any of them. The table is not the argument; the sources in it are.

Post-market surveillance, PMCF and vigilance

Three questions about the same discipline: making a post-market claim that can be reproduced by someone who was not there.

7. What is a realistic post-market trending threshold for rare events?

Asked live: events not reasonably quantifiable in a clinical study, with no device-specific investigation data. The asker proposed a threshold on the increase in occurrence to trigger action, plus an absolute occurrence threshold to trigger notification.

Short answer: the two-tier structure proposed is the right shape. What makes it defensible is naming the denominator and the statistical rule in the PMS plan, before the period starts.

Trend reporting is Article 88: a statistically significant increase in the frequency or severity of incidents that are not serious, or of expected undesirable side-effects. Article 88 does not leave the method open. It requires that “the manufacturer shall specify how to manage the incidents referred to in the first subparagraph and the methodology used for determining any statistically significant increase in the frequency or severity of such incidents, as well as the observation period, in the post-market surveillance plan referred to in Article 84”. Annex III, section 1(b) reinforces it, requiring “suitable indicators and threshold values” in that plan. So the rule and the period are prescribed to be written in advance, which is the whole answer to this question.

For rare events, an absolute count on a small denominator generates false alarms and then gets quietly ignored, which is worse than having no threshold. Four things make the rule hold:

  • A stated denominator. Devices sold, units implanted, or patient-years, chosen once and written down. Most trending arguments fail here rather than at the threshold.
  • A rate, not a count, so that a growing installed base does not read as a safety signal.
  • A statistical rule chosen in advance, before the data exists. A control-chart rule or a comparison against an expected rate both work; what matters is that the rule is fixed beforehand.
  • A separate, lower notification threshold keyed to severity, where a single occurrence is enough to act. That is exactly the second tier the question proposed, and it is the right instinct: frequency rules and severity rules should not share a threshold.

And write down what happens when the denominator is too small to trend at all, which is the honest position for genuinely rare events. A plan that says so, and names the alternative signal it will use instead, is stronger than one asserting a threshold that could never fire.

8. Does the CER need its own vigilance search, or is internal complaint data enough?

Asked live: the MDR requires vigilance searches but does not spell out subject device against similar devices. Is internal complaint, FSCA and CAPA data sufficient, given manufacturers are notified of adverse events and recalls through public databases?

Short answer: internal data alone is not sufficient, because it cannot tell you what happened to comparable devices on the market.

Your complaint, FSCA and CAPA records cover your device inside your own system. They are necessary and they are not the whole picture, and the Regulation says so directly. Annex III, section 1(a) requires the post-market surveillance plan to collect “relevant specialist or technical literature, databases and/or registers” and “publicly available information about similar medical devices”. Section 1(b) requires a proactive and systematic process that “shall also allow a comparison to be made between the device and similar products available on the market”. Neither is satisfiable from your own complaint file, and the clinical evaluation draws on that same post-market data.

The practical requirement is reproducibility. Run the search the way you run a literature search: a defined list of databases, the exact search terms, the date the search was run, the period covered, and the screening applied to what came back, all recorded in the search protocol. Someone should be able to re-run it and land in the same place, which is the same test that applies to the literature search.

One trap to name in the report: public adverse-event databases have no denominator. They tell you what has been reported, never an incidence rate. Counts from them are signals to investigate, not rates to compare against a threshold.

9. Is one year of adverse-event database data acceptable for a PMCF evaluation report?

Asked live: the submission falls mid-year, and the database records are not static, they are updated through the year.

Short answer: yes, if the window is defined and justified. The problem is almost never the length. It is an undated extract.

State four things and a one-year window is defensible: the exact database and interface used, the exact query, the date the extract was taken, and the period the records cover. Because records are added and amended continuously, the extract date is part of the result. The same query re-run three months later will not produce the same table, and if the date is missing, the discrepancy looks like an error rather than the expected behaviour of the source.

Justify the window against the reporting cycle rather than convenience: align it with the PMS reporting period so the PMCF evaluation report, the PSUR and the CER update are all reading the same interval. A window chosen to fit a submission date, and no other reason, is the version that draws a question.

Then state the limitations of the source in the report itself: voluntary and inconsistent reporting, no denominator, duplicate entries for the same event, and coverage of one market only. Naming them is what stops an absolute count being read as an incidence rate later.

Clinical development and investigation strategy

Two questions about the document that is most often written last and should have been written first.

10. Clinical development plan: what changes between a new device and an established one?

Asked live.

Short answer: for a new device it is a forward plan. For an established device it is mostly a map of what already exists, plus what is still open.

Annex XIV Part A asks for a clinical development plan showing the progression from exploratory investigations through to confirmatory investigations and on to PMCF, with milestones and a description of potential acceptance criteria. That is the same requirement in both cases; what differs is where the device sits on that path.

A new device: the plan is genuinely prospective. Feasibility or first-in-human work, then the pivotal investigation, then the post-market programme, each with the question it answers and the criteria it will be judged against. The value of writing it early is that it forces the acceptance criteria to be set before there is any data to be influenced by.

An established or legacy device: most of the path is behind you, so the plan documents where the evidence already sits, what each certification cycle produced, and which questions remain unanswered. This is also where continuity has to be visible: if a threshold has changed since the last cycle, name the change and justify it against published evidence rather than letting the reader find the difference.

The failure mode is identical in both cases. A clinical development plan with no dates, no acceptance criteria and nobody accountable for the next step is not a plan, and it reads as one written to fill a required heading.

11. Can clinical data from outside the EU alone justify not running a pre-market investigation?

Asked live: what specifically would a reviewer look at to decide whether foreign data is sufficient and applicable to the target population, for safety, performance and clinical benefit?

Short answer: the origin of the data is not what decides it. Whether an investigation can be omitted is decided by Article 61, and applicability is decided by how the differences are handled.

Two separate questions get merged here, and separating them is most of the answer.

Whether an investigation is required at all is answered by Article 61. For implantable and class III devices, a clinical investigation is required unless one of the narrow exceptions in Article 61(4) to 61(6) applies. Foreign clinical data is not one of those exceptions. It cannot, on its own, remove a requirement that the Regulation puts in place.

Whether foreign data is usable is a separate matter, and it usually is usable. The MDR sets no geographic condition at all: Article 2(48) defines clinical data without any restriction on where it was generated, and the Regulation carries no provision on data generated outside the Union. The conditions come from guidance. MEDDEV 2.7/1 Rev. 4, Appendix A1 asks for a justification that “should explain if the clinical data is transferrable to the European population, and an analysis of any gaps to good clinical practices (such as ISO 14155) and relevant harmonised standards”. Applicability then turns on named differences:

  • Population. Demographics, comorbidity profile, disease severity and stage, anatomy where it matters.
  • Standard of care. What the comparator arm actually received, and whether that is what a European patient would receive today.
  • Device identity. Whether the configuration studied is the configuration being CE marked, including software version, sizes and accessories.
  • Endpoints and their measurement. The same instrument, the same time points, the same definition of the event.
  • Conduct. Whether the study was run to a standard equivalent to ISO 14155, with the monitoring and data integrity that implies.

What decides the outcome is not the country. It is whether each of those differences is named, its impact assessed, and the residual uncertainty carried into the PMCF plan. A file that lists the differences and closes them is in a stronger position than one asserting the populations are comparable and moving on.

Still open

Forty-seven people submitted a question when they registered, across seven themes, and the hour covered a handful of them. The eleven above are the ones written during the session and never reached. The registration questions are being answered in the same format and will be added to this page.

If your question is not here, or the answer raises the next one, put it in the comments under the video. They get answered.

Watch the Full Session →

Continue the Conversation

Join the Clinical Evaluation Navigator community for weekly expert calls, and work through real clinical evaluation problems with regulatory professionals facing the same ones.

Join the Community

Free to join • Weekly Friday calls • 100+ published articles