Showing posts with label assessment. Show all posts
Showing posts with label assessment. Show all posts

Friday, March 4, 2016

Using a Lego to explain the difference between competencies and EPA's

People in medical education often have trouble figuring out the difference between competencies and EPA (entrustable professional activities). There is a pretty big philosophical difference. The competencies are definitions of observable behaviors and the EPA's are about observing a learner do a specific work task. Here is a recent article from Carraccio and others  that tries to ties the concepts together.

I was in a meeting yesterday where we were discussing the differences between EPA's and competencies. The group was trying to determine whether you are obligated to assess one first. We have 43 competencies in our new curriculum and 13 EPA's. The question that came up was if EPA 2 is going to be assessed in a student, and it is identified to require multiple competencies, do I need to measure the competencies first to allow me to get into an assessment for EPA's? The reverse of this question is if I am found to be entrustable to an acceptable level for graduation for EPA2, does this automatically allow me to be entrustable on all the related competencies.

While this discussion was going on, my mind wandered to Legos. I've been building Lego sets for years. My son and daughters now have large tubs of Legos in our house. It's really cool how you can make all sorts of wonderful things with the simple building blocks that are Legos. You can think of competencies as the individual building blocks. These are the behaviors necessary to build cool stuff. If you don't have the basic building blocks, you can't really make many cool sets (EPA's). The blocks come in lots of different shapes. Think of each shape as a competency. There are long flat short pieces and long flat long pieces. There are two by four bricks and two by eight bricks. There are all sorts of bricks. The bricks also come in different colors which can represent that a competency must be demonstrated in many different environments prior to saying for sure that it has been acheived. In other word, you may be good at applying medical knowledge in a pediatrics outpatient clinic, but not in an inpatient ICU with a critically ill patient. So, to check out on any given competency, the student may need a green two by four brick (applying medical knowledge in peds clinic), and a red two by four brick (applying medical knowledge in an ICU).

EPA's then are like the ability to build the sets. An EPA would be like taking all those Lego bricks and putting them together to make a car or a boat or a house. The act of making the car or boat or house means that you not only have the bricks needed to make the set, you can use them appropriately. So to enter an order in the ICU would be like making a house. The learner needs some two by four red bricks to make the house, but will also need roof pieces (say an infomatics competency) as well as other pieces. And they need to all be the right color to make a house in the ICU setting. Having a red two by four brick does not mean a student can build a house (they need specific skills to put it all together and other pieces), and building a house in the ICU does not mean you can build a house in the peds clinic (you need green pieces for that).

So, in other words, the EPA's and competencies are each dependent on each other. But both need to be assessed in parallel to assure that students Lego buckets are full of lots of cool and useful pieces, but also to assure that they can actually use the cool pieces to make stuff. Let me know if this helps you understand how EPA's and competencies work together, and what you think of this analogy in the comments below.

Friday, July 19, 2013

Assessment in the new world of CME

I'm at a workshop at the American Academy of Neurology headquarters in Minneapolis working on an online CME activity I am creating for their Neurolearn series.  The AAN is working to create online educational tools for neurologists (although other healthcare providers are welcome to view them as well).  It's my first experience with creation of online, interactive CME.  I'll admit the interactive portions of these courses is still pretty rudimentary, but it got me thinking about how to design these courses as the learning environments become more complex.

What I've come to realize is that in the CME world of the near future, assessment is going to be huge.  If you truly want to develop the learner centered environments, you need to have very solid assessments.  The reason is that if you have a flexible learning path on a particular topic (say fall prevention which is the topic I'm working on), you need to quickly and deftly sort out the novice to expert continuum on the topic within a very brief amount of time.  If you are relegated to MCQ, you probably will have about 5, and at most 10 questions the learner will likely tolerate before you have to shunt them into their learning tract.  And if you mis-align folks and put them in the wrong tract, that will also be potentially disastrous.  I would define disaster here as the learner aborting the course before they have completed it.  A secondary disaster would be someone who is so bored or confused by the content that they complete it, but do not really pay attention to what was presented.  So, your 5 questions need to be very focused, and very strongly crafted to allow you to sort novice to expert quickly.  That will take a lot of effort to achieve that.

So, future CME creators of the world.  I'd strongly encourage you all to consider taking advanced courses on assessment before learning how to do cool stuff through web design.  Just my thought.

If you want to check out the current NeuroLearn courses which are available look here.

Friday, August 10, 2012

Clinical assessment variability - what is really causing it?

There was a recent article in Academic Medicine by Dr. Alexander and colleagues from Brigham and Women's Hospital describing the amount of variability in clerkship grading among US medical schools.  They found that, unsurprisingly, the grading systems for the clinical years had really no consistency at all.  There was inconsistency among the grading systems used (traditional ABCDF or honor/pass/fail or pass/fail) - (table 1), and even within the schools which used a similar scale the percentage of students receiving the highest grade was all over the place (table 2).  So, the question is what do we do with this information?  I think no one really expected findings that were different, but now the answer is out there, in print (or on digital reader screens).

I think part of the answer to where we go from here is to decide if this article was really asking the right question.  The authors do start to talk about this in the discussion section, but I'll try to lay out my thoughts with a little different spin than they gave their discussion.  I think the real question is what are we using the assessment of the clerkship performance for?  What is the essence of what we are trying to measure?  Only when there is broad consensus not only between schools, but within the individual courses of each school will there get to be any semblance of uniformity of grading of students.  I see at least two competing interest which influence how a clerkship director decides to come up with a grading system.  The first is the idea that the students should be measured on how competent they are in the area the clerkship is grading.  In other words, when they are on call as a first-year resident or as a 50 year-old physician, do they have the knowledge and skills to assess a patient with a given problem.  Second, the clerkship director also wants to be sure that the students at their school have a fair chance to compete for selective residency programs.  Thus, there also needs to be a system to distinguish high-achieving from low-achieving students.  The first system is more about the individual student, and with this system, by definition, everyone should be able to achieve the highest score with enough effort and work.  In the second system, it is more about evaluation of the program, and the group.  In this system, it cannot be possible for everyone to achieve the highest score.  However, the system can be manipulated on both sides to aid students or to make it more hazardous.  There are benefits and risks of each system - as with anything in medicine.

I don't think these interests are necessarily incompatible, but they create a tension which I've seen in national meetings and in local curricular meetings.  I also think most clerkship directors are not aware of how this tension affects the grading system they have developed.  I think their not aware as the debates I've heard are usually about tools for assessment or the numbers of honors.  Rarely does the debate get to the level of what is our ultimate purpose for the assessment.  The answer to that question must shape how grades are assessed.  Only when we all become very clear about what we our goals are for the assessment will we truly be able to come to a place where we can have a national dialogue about how to unify the system.

Wednesday, June 6, 2012

EBM evaluation tools applied to medical student assessment tools

I remember back to the days when I was a fresh medical student taking those first classes in biochem, anatomy, and cell biology.  I learned a ton, and honestly I draw on this knowledge-base daily when I'm taking care of patients.  I also remember that the assessments methods used during my first year of medical school were not the greatest (in the opinion of a person who was teaching high school physics and chemistry 3 months before entering med school).  The number of assessments used in med schools has risen over the last 15 years since I was an M1.  However, with a rise in number of choices, comes responsibility to utilize the right choice.  Another way to look at this from an pedagogical standpoint is are the assessments really measuring the outcomes you think they are measuring.  To attempt to help the medical educator with this dilemma, I came up with the idea that you can apply a well-known paradigm used to evaluate evidence-based medicine (EBM) to evaluate a student assessment.  The EBM evaluation methods I've been most familiar with is outlined by Straus and colleagues in their book, Evidence-Based Medicine: How to Practice and Teach EBM, copyright 2005.

Here's my proposed way to assess assessment:

1)  Is the assessment tool valid?  By this we need to be sure that our measurement tool is reliable and accurate in being able to measure what we want it to measure.  The standardized (high-stakes) examinations like MCAT, USMLE and board certification examinations are expensive not because these companies are rolling in cash, but because it takes people LOTS of time to validate a test.  Hence, most home-grown tools are not completely validated (although some have been).  To be validated an assessment has to be likely to give similar results if the same learner takes the test each time.  It also has to accurately categorize the level of proficiency of the learner at the task you are measuring.

For example, let's say I have an OSCE to assess whether a learner can counsel a young woman of child-bearing age on her options for migraine prophylaxitic medications.  For my OSCE to be valid, I need to look for reliability and accuracy.  Does the OSCE predictably identify learners who do not understand that valproate has teratogenic potential, and don't discuss this with a standardized patient?  You also want to know if it is accurate, in other words does your scoring method give similar results if multiple faculty who have been trained on how to use the tool score the same student interaction?  To truly answer these questions on an assessment, it takes multiple data points for both raters and learners - hence why it takes time and money, and also why most assessments are not truly validated.

The best way to validate is to measure the assessment against another 'gold standard' assessment.  How well does your assessment work compared with known validated scales.  Unfortunately, there aren't as many 'gold standard' assessments outside of the clinical knowledge domain in medical education (although it is getting better).

2)  Is the valid assessment tool important?  Here we need to talk about whether the difference seen in the assessment is actually a real difference.  How big is the gap between those who just passed without trouble, just barely passed, and those who failed to meet the expected mark?  Medical students are all very bright, and sometimes the difference between the very top and the middle is not that great a margin (even if it looks like it on the measures that we are using).  I think the place where we trip up here sometimes is in assuming that Likert scale numbers have a linear relationship.  Is a step-wise difference from 3 to 4 t o 5 on the scale set up on the clinical evaluations a reasonable assumption, and is the difference between a 4 and a 5 really important?  It might very well be that this is true, but it will be different for every scale that we set up.  I've never been a big fan of using Likert rating scores to directly come up with a percentage point score unless you can prove to me through your distribution numbers that it is working.

3)  Is this valid, important tool able to be applied to my learners?  I think this step involves several steps.  First, are you actually measuring what you'd like to measure?  A valid, reliable tool for measuring knowledge (typical MCQ test) unless it is very artfully crafted will not likely assess clinical reasoning skills or problem-solving.  So, if your objective is to teach the learner how to identify 'red flags' in a headache patient history, is that validated MCQ the best assessment tool to use?  Is it OK that that learner can pick out 'red flags' from a list of distractors, or is it a different skill set to be able to identify this in a clinical setting?  I'm not saying MCQ's can never be used in this situation, you just have to think about it first.

Second, if you are utilizng a tool from another source and you did not design it for your particular curriculum, is the tool useful for the unique objectives?  Most of the time this is OK, and cross-fertilization of educational tools is necessary due to the time and effort bit.  But, you have to think about what you are actually doing.  In our example of the headache OSCE, let's say you found a colleague at another institution who has an OSCE set up to assess communication of differential diagnosis and evaluation to a person with migraine who is worried they have a brain tumor.  You then apply that to your clerkship, but you are more interested in the above scenario about choice of therapy.  Will the tool still work when you tweak it?  It may or may not, and you just need to be careful.

Hopefully you've survived to read through to the end of this post.  Hopefully you learned something about assessment in medical education, and you found the EBM-esque approach to assessment evaluation useful.  My concern is that in general, not enough time is spent considering these questions, and more time is spent on developing the content then on assessment.  I'm guilty of this as well, but I'm trying to get better.  Thanks for reading, and feel free to post comments/thoughts below.

Tuesday, March 6, 2012

What medical education can learn from "Moneyball"

I've been waiting a bit to write this post, as I'm not sure exactly which way to take it.  Let me start by stating that I'm a really big baseball fan, and have been since second grade when my dad first took me on the El in Chicago to see the Cubs play in Wrigley.  I still get chills walking into that place.  This love of baseball drives me read the occasional baseball book.  So, while I haven't seen the recent movie, I read the Michael Lewis book, "Moneyball," a few years ago.  And I really liked it on many levels.

In the realm of medical education, I liked the idea of trying to measure something that is inherently immeasurable.  In some respects, trying to pick a good candidate from a pool of medical school applicants or trying to assign a grade to a student on a clinical rotation is not unlike what the old-time scouts in "Moneyball" were doing.  They would look at a player batting, pitching, or fielding, and go with an overall geschalt of whether that player was 'big-league material'.  They were also basing their decisions on statistics which had been around forever, and no one had ever really questioned whether they worked or not to predict who is or who is not going to be a good performer.

Then, Billy Beane and his team of statisticians looked beyond the traditional numbers and redefined what to look for in a player prospect by largely ignoring the players current body habitus or mechanics and focusing solely on the numbers.  They also redefined what success was by finding the the number of runners on base per game correlated to wins more tightly than other statistics.  Thus, on-base percentage, and slugging percentage (which measures walks with extra-base hits) was more important for how an individual would contribute to the team than total runs batted in or home runs.  (Sorry if I just lost the non-baseball fans out there).

This process can have applications to lots of venues.  I think medical school needs to re-look at how we are evaluating our students and decide if we need to go through a similar process.  Are there statistics available to us now which may not have been available 20 to 30 years ago that we could use to identify medical students who are not likely to do well in practice.  We're pretty solid at identifying people with knowledge gaps as our system of standardized testing takes care of that.  But, is that what really makes a good physician?  It's a part of it for sure, but it is not all of it.  There's a lot more to clinical reasoning, and professionalism than just knowledge base.  Can we find ways of identifying ways to capture those measures, or are we going to be stuck with the old scouting reports and crossing our fingers to see what happens?  I don't have any solid answers yet, but I'm willing to help look.