tech_surveillance1745 wordsRead on Arc Codex

Starting a Career in Data Science in the Age of AI

Starting a Career in Data Science in the Age of AI How do you set yourself up for a career that will stand the test of time when things are changing so fast? I was recently asked by a college student in data science and computer science for my advice on entering and succeeding in the field of data science and machine learning today, without selling your soul or sacrificing your ethics. It was a really hard question to answer, because I myself entered the workforce almost 20 years ago, and the field of data science 10 years ago, and so much has changed since that time. Questions of ethics were much less pronounced (although not absent) when I became a data scientist, and today they are front and center. I did my best to answer in the moment, but I’ve spent some more time thinking about it, and I feel like there are a few key considerations for students trying to map a path to a rewarding career in our field that will last. Diversify Industries First, working in the software industry is not necessarily the path for all of us. Data science broadly, and machine learning engineering in particular, have fantastic applications in pretty much every sector, much more so than in the past, so I encourage students to consider fields like healthcare, government, nonprofits, and other sectors. Data can make a massive impact on the success and efficiency of all kinds of work when analyzed well and modeled in a sophisticated way. Internships are a good way to find out what data scientists in different sectors and industries do, and I encourage students to just simply read the job descriptions as well, and talk to practitioners when possible, to help get a feel for the expectations. The character of the roles will change (often rapidly) but there’s no better way to figure out what the job can be like than to ask people who are doing it. Data Education On the topic of ethics in the practice of data science, I believe it’s the responsibility of data scientists to educate colleagues about ethical, effective data utilization. I’ve written about this many times, but we are the people in prime position to help whole organizations understand how to use data and models, including LLMs, in ways that produce results and don’t endanger people’s safety, health, privacy, or welfare. It’s incredibly, horrifyingly easy for organizations without proper data expertise to careen ahead with AI and other technologies with ignorance of the risks, and it should be part of the practitioner’s role to guard against that, just like it’s an accountant’s job to prevent the company from committing financial crimes. If you want to feel good about your work and sleep well at night, taking this responsibility seriously and being a voice for ethical, safe applications of machine learning is a good place to start. Solving Problems What else should you expect to do in a job as a data scientist or MLE? Well, expect to be solving problems. That’s really where data science shines, where business problems or uncertainties appear, and knowing the answer is vital to making a correct decision. You won’t start out finding your own problems, instead your manager and leadership will be defining and scoping the questions that need answering, but in a good role, you’ll have some autonomy in choosing the strategy and methods you use to answer. By doing this, you’ll learn how to approach different kinds of data questions, and you’ll see how the problems are defined and where they come from, and this kind of experience is invaluable. That’s what differentiates entry level from senior level practitioners in data — the ability to recognize a problem’s archetype, hunt through messy and complicated data to develop a plan, and then to execute on it, generating a statistically rigorous and sound answer. This skill set, in my experience, can take you to most any industry or sector that interests you, because the real skill is problem solving. You’ll learn a lot in college (if you do it right) and you’ll need that knowledge to progress, but you can’t learn the real problem solving skills until you’re out there in the field doing it, in my experience. Don’t Over-Specialize In contrast, training and building frontier models is not the way to a lasting career for most. Most of us working in the field will never get near “cutting-edge” frontier model construction, and I for one wouldn’t want to. Technologies change, and methods of performing machine learning tasks change — I have witnessed this firsthand. Some narrow specialists will focus on training LLMs, but I recommend flexibility and versatility, especially in your early career, so that you have more options as the field inevitably changes around you. I am glad that I have a good foundational knowledge of neural networks and can train and tune them, but I am equally glad that I also know GBMs, NLP, clustering, and other techniques, because these are just all tools in my toolbox, which I can use when they’re appropriate. LLMs are not going to be the end of machine learning, and we need to be watching for whatever comes next. Use LLMs Cautiously You might also be wondering how and when you should expect to be using LLMs in your professional life, and this is a question most everyone in white collar fields is asking, not just data scientists. I struggle with this, because AI is a useful tool in many situations — writing code, for example — but abdicating our critical thinking processes to the chat bot is dangerous. I’ve talked about this elsewhere, including discussing the financial implications, and I encourage entry level folks, at least for now, to make sure they have the capability do the job manually, even if they don’t have to every day. I don’t recommend taking a rigid stance of refusing to use it, because in coding at least you really may fall behind peers, but I propose finding ways to use it for what it’s good for, and being selective. I routinely get complimented at work because I write all my own text (like these articles, where I never, ever let AI touch them) because it sounds and feels human. People appreciate the fact that I care enough about the work and about my readership to do the writing myself. Furthermore, it’s rewarding to know that my own brain and two hands produced this work without any augmentation from AI. (It’s also rewarding when I write my own code and don’t get AI help, but I compromise in the service of getting the job done, because I’m not necessarily convinced my code would be better than what an LLM would produce with my guidance. I am, however, absolutely sure that my writing is better than what an LLM would produce, because I have faith in my ideas and my own creativity.) The student who originally asked me this question talked about feeling quite frustrated with peers who completely rely on AI for schoolwork, and never make the distinction between tasks where the journey is the destination, so to speak. I think this is a common concern for young people, and I hope the anti-AI backlash that many young people are displaying will curb it, but it remains to be seen. Anytime you (or I) are doing a task where the objective is to LEARN, AI shouldn’t be involved in performing the task. Practicing and performing the task is the best way to learn the skill. In the workplace, there are still some tasks where learning is a key part of the objective, but often the goal is more around completing the work. That difference is integral to understand, to know when using an LLM might be appropriate and when it’s not. Ask yourself, 1. Is the point of this, or one of the goals, that I learn something? and 2. Is there any difference in the quality, character, or nature of the work that means I can do it better than an LLM? and if the answer to either one is Yes, then you need to do it yourself. Expect the Unexpected This is all to say that setting out on a career in data science is fraught with many unknowns. This field has changed profoundly in the past decade, and we should expect that to continue. The next big idea in machine learning is coming, but we don’t know what it is yet. We just know that practitioners are going to be expected to learn it and figure out how to apply it to business problems, just like we have had to do with the other technologies that have come on to the scene. I can’t promise you “study A, B, and C, and you’ll be able to get a good job and build a career that will carry you to retirement” because that’s just not the world we live in. Global economic events will come along that will make finding work easier or harder, and technological or political phenomena will shape the economy so that certain skills are more valuable than others. Finding ways to be valuable by being able to solve problems, using whatever technology comes along, is the best advice I have for maintaining relevance and employability in the extremely unpredictable future. Questioning the Premise Some readers might wonder why I don’t challenge the premise, and argue against the economic system and society that forces us to find a way to be valuable and employable to justify our existence. It’s tempting! I do have strong feelings about the economy’s structure and the cruelty of the system in which we live. But questioning the system, while valid, doesn’t help the young people who still have to live within it right now. A college student, especially one without significant resources or structural privilege, needs to find a way to pay the rent and buy groceries, before they can be in a position to effectively challenge capitalist ideologies. So, I am answering the question of how to get there, to the best of my ability. I hope that college students and junior data practitioners won’t take this as any kind of endorsement of the system, because it isn’t one — it’s just an acknowledgment of the realities we’re all facing today. Read more of my work at www.stephaniekirmer.com. Further Reading https://www.stephaniekirmer.com/writing/machinelearningspublicperceptionproblem/ https://www.stephaniekirmer.com/writing/closingthegapbetweenmachinelearningandbusiness/ https://www.stephaniekirmer.com/writing/thenewexperienceofcodingwithai/

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.