I’ve tried to write several times before, on my blog, on medium, on Quora, and each time I got pretty good audience responses—being cited in conference talks, getting hundreds of messages from internship-seeking students1, and most recently my long-lost school friend from 5th grade reconnecting ~18 years later and telling me he found my writing funny and I should write more (he is to blame for giving me the impetus to produce some of this bespoke human slop).
This article is about my childhood dream. I felt it only right to (re)start my writing journey here.
I’ve always wanted to change the world. Ever since I was in 8th grade, and you asked me what I wanted to do, I would say “I want to change the world.” Now I didn’t know what it means to change the world, I certainly didn’t know how I’d do it, and it was arguably more a tongue-in-cheek response than an actual goal. But I had the chutzpah2 to say it out loud, and not without intent. And I think that has mattered the most in my career and life choices. This article will be part reflection, part reminiscing, perhaps part hubris, but entirely true. I think it’s a useful way to share who I am and why I do what I do, since I get asked that question once a month now.
Believe it or not, I don’t think in the years that followed since 8th grade, that this life goal changed even once. Now it wasn’t actively on my mind, like I wasn’t directing every action towards this end because I really didn’t have any idea of what the process is for this world-changing to happen. But I knew that I would have to start by fixing problems I saw in front of me even if nobody else cared about them. Because technically, that problem was part of the world, and technically, fixing that tiny problem did fit the criteria of changing the world. But more practically I just wanted to keep learning new things and hoped it would help explain how the world worked so I could figure out what I wanted to change about it. I worked in particle physics on tracking jets (“quantum stuff”), built and shipped content recommendation algorithms at terabyte scale, designed online safety systems uncovering state-run influence operations, built digital health literacy platforms deployed in small towns in India and Bangladesh, and built one of the first agentic marketplaces to run behavioral experiments. But it was all in service of the same goal—to figure out how the world works. And in the process of exploring these really scattered areas, I noticed that often when you get too deep into an area, you gain expertise not necessarily because you’re so good at it, but because noone else cares about it enough. And if you pick problems right (or just get lucky), and they later become important, suddenly you’re the world-leading expert in this area everyone wants to work in, but doesn’t have the expertise to do so.
Problems aren’t fixed not because people don’t notice them but because most of us don’t care enough to ask twice why they exist.
On the other hand I was simultaneously fielding emails from one of the directors of my program that they had tried and failed to find me funding or another advisor with a not-so-gentle suggestion that I could “Master out”. It was a time I would not like to relive, but things ended better than the email foresaw and several years later the same person would invite me to a coffee chat because they found my work super interesting. Life comes around, but let’s stay on story.
I have been a naturally curious kid and that was one trait that still drives my problem selection3. I spent some time exploring very disjoint ideas like hardening BSD-based operating systems by stripping down microservices, integrating e-commerce analytics into consumer Android applications, developing surrogate models to approximate complex physics simulations, scaling probabilistic inference and causal discovery. I enjoyed each of these problems and spent 6-8 months intensively exploring each one in internships, research stints, and software engineering roles. But it didn’t feel like I was drawn to them deeply enough to spend all my waking hours understanding each problem—and to be honest I felt like I could never be the best in the world at them. And then came an opportunity to work on social media and politics which changed my career and got me a step closer to that childhood dream4.
In the second year of my Data Science Ph.D., in an internship at Adobe Research5 I found myself working on content recommendation algorithms and how they shape our feeds. I collected terabytes of data on my local device which itself had 500 GB of storage, figuring out data acquisition, streaming, indexing, embedding, and learning, figuring out how to collect public content from some popular platforms. This experience really stuck with me because first, I realized even as a major multimedia-product organization, Adobe lacked some multimodal infrastructure I would have considered pretty fundamental to actually build production-grade content recommendation algorithms. Now don’t get me wrong, they had a ton of raw materials like compute, data, labels, and models, but just lacked the right set of pieces chained together that would really have helped me out. I ended up using open-source tools and architecting my custom solution from scratch for the specific problem I was working on—and surprisingly that made me the one of the most experienced people at the company on tasks of the kind I published6. So I realized while it is theoretically correct that “FAANG could disrupt any startup given their massive resources”, in practice FAANG would rather acquire the builders with existing credibility and give them the raw materials to build it. This way they eliminate the risk of wasting their substantial raw materials building the wrong solution—and it’s just way more efficient (cost, time, and risk-wise). So the acqui-hiring spree in the AI space made complete sense to me (the valuations though, are debatable).
Following Adobe, I restarted my pending work at NYU CSMAP, where I was now based as one of the Data Science Ph.D. students they adopted kindly, and introduced to political science. I learnt about text-as-data methods, took a course on causal inference, and started to pursue a new line of research exploring the causal effects of content moderation interventions. Specifically I wanted to understand if interventions to limit misleading information deployed on one social media platform could impact the information that spreads on another social media platform. This was super interesting to me, because platforms don’t study each other, certainly rarely look at cross-platform information let alone interventions. Yet as consumers we are still exposed to the same harms on multiple different social media platforms. And despite the recurring harms, we have no agency to try and limit them as consumers. So it fell to a researcher to start to painstakingly collect information across multiple platforms and organize it for everyone to have transparency into cross-platform information spreading. Then, we might have some understanding of interventional impact across platforms. This was one of the first papers of its type and I was super excited to see the results which indicated that interventions could actually be effective in mitigating cross-platform harms not just on their own platform! But during this project I realized there was a potentially bigger problem to solve. Noone seemed to have great ways to collect cross-platform online posts! Even though I did it using existing tools provided by platforms, I was not sure why a platform would even provide them? What’s in their interest here? Why would a platform ever invest into providing researchers with data access from their system when the most popular papers appeared to lean towards investigating the harms on social media platforms more than the benefits thereof7. It was a fortuitous feeling because in just a month or two more, X had announced an API shutdown and released pricing terms for their API. By that time I was already architecting my own infrastructure to build platform-free methods to collect their data — I found legal ways to collect public data from social media platforms without needing to go to them to access it. I was one person working on it early days as this was an incredibly exciting idea that nobody really cared much about. I tried hiring interns and that was a terrible mistake. But eventually I built the first version of the the platform that I had wanted to build. Turns out I also got a grant from the Wikimedia Foundation to build Arbiter, a cross-platform tool to investigate coordinated networks of accounts. This article was published in early 2023, though my grant was in mid-2022. The first round of tech layoffs hadn’t happened yet. Trust and safety was still an active priority. ChatGPT hadn’t launched. And I made a bet that this is the problem I will be the best in the world at solving, and a problem that’s important enough to solve regardless of what else happens in the world around us.
That feeling of one person being able to drive big changes really solidified in my next internship at Twitter8 where I spent a literal 11 weeks out of 12 weeks of my Ph.D. Machine Learning Engineering internship just looking into data9. So much so that I recall the then-CEO actually ended up chatting with me for two whole minutes in our intern happy hour10 asking if I had the data that I needed to do my job. I did not. In week 11, I was gently told I wasn’t getting a full-time job offer, based on my work. In that last week, as planned, I ended up building and deploying a series of iteratively improving classifiers that ultimately maintained the recall performance and beat the precision of the existing classifier for civic misinformation by a lot. The really cool part was that my classifier was an early detection system, so unlike with the existing classifier, Twitter wouldn’t need to wait until information spent too much time on the platform and it could still perform better than the existing classifier which only worked once content had spent some time gathering engagement on the platform! Now I don’t know whether the team got laid off before or after deploying it because human labeling was still ongoing and would be ideal to have before deploying any new models, but what I do know for sure is I ended up getting that return interview offer in the end. But I realized something really important.
Everything that I built at both platforms was hosted on non-proprietary systems. With the exception of the data, everything I built at Twitter could be built on the outside by someone with the right combination of novel ML research ideas and software engineering skills. It’s just that nobody seemed to care enough about this problem to try to solve it properly. So I decided that someone would be me. This was how I would change the world.
You know the scene in Kung-Fu Panda where Po the Panda opens up this golden Dragon Scroll and sees nothing written in it and then realizes the answer is him? There is no secret ingredient, it’s just him? Well this was kind of my reaction because while the movie plays it out nicely and says he’s the hero of this story, what people don’t realize is this also means there is no backup plan. It’s just him, and if he fails then everything goes to shit. So this is the appropriate response a human being should have when they realize this is how it is, when you solve problems other people haven’t bothered to solve yet.
And that checks out because this is also how startups emerge, if you’re considering the organic ones11. There’s a need created because existing problems are overlooked or perceived as either too hard or too trivial to be worthy of solving. Most startups fail, and that’s a really scary prospect—to see what you’ve built crumble before your eyes. You are the final straw, the last one standing, the end of the line, the buck stopper, what have you. And that prospect is consistently scary as shit for all kinds of people doing a startup. So Po is enacting an accurate feeling, actually. But at the same time there is no better feeling than waking up tired-yet-happy from working incredibly hard on a problem that you are satisfied to dedicate your life to. And that’s exactly why—despite the high likelihood of failure—it’s still worth it to try.
I mentioned I was always curious as a kid but that doesn’t just mean I was exploring problems. I was just as actively thinking through prospective solutions to the problems that I saw in the world around me. And I discovered the receipts to show for it. Every since 2012-13, I have been writing down tens of ideas in a journal, later turning into blog posts instead of just ideas. A few years ago, I realized I still had this journal in my saved memorabilia and I opened it up just to see how my ideas were 11 years ago. And I was pleasantly surprised to find past me was not as much of an idiot as I’d expected—or at least I’d remained the same idiot over the years without changing my vision for the future. And one idea in particular would prove this beyond a reasonable doubt—SimPPL.





Turns out, past me had pretty much predicted exactly what my future nonprofit, SimPPL, would start out doing almost 7 years later. I spelled it out, that we would try to offer internships and student projects of all kinds from gaming to web development. And that’s exactly how I designed our Google and Mozilla-awarded NextGenAI program. We structured it just like student internships might look like, putting teams through the wringer to define and iterate on a value proposition, and eventually pitch their ideas to get funding—with four receiving commercialization offers at the end of the program! We took this program’s lessons to inform the NYU AI School that I led and then advised from 2019-2024. And my cofounder ended up taking the program to Germany at the University of Mannheim, with one of Germany’s best business schools.
This educational programming was the early work done over ~3 years that formed the basis of our tech nonprofit. We ran these programs and saved some of the money to do what we really wanted to do: build our AI platform, Arbiter, to combat public deception campaigns and protect civil discourse. Yes, that same Arbiter from 2022-23. I was still holding out because I believed in the value of that problem. And eventually we did build a solution for it that worked. Today, Arbiter operates in 23 countries and counting, with 150-ish newsrooms, civil society orgs., and even regulators signed up to use it. We are invited by a state government in India and a state government in the U.S. for a statewide rollout of the platform, and two other state governments are reaching out! Arbiter helped us surface prospective job scams, false electoral claims, harassment of minorities, influence of the manosphere on younger audiences, combat child trafficking, and more (some of the data collected for these studies are here).
With Arbiter, we can show you what shapes your feed before it shapes your views.
Our work led to tangible outcomes like national news features, causing the removal of harmful networks harassing political actors by a major social network, invitations to advise policymakers, directly influencing regulation, advising heads of state and world leaders, and tens of invited talks to industry and academia. Last week, someone on the flight sitting next to me saw SimPPL’s work and made us pitch to their neighbor who got us into a room where Paul Graham was giving a talk to the current YC cohort—we were the only nonprofit in the room, I suspect. Last month, Stanford selected us to host an invited workshop in October a their flagship trust and safety research conference for showcasing Arbiter. This was a slot previously held by Meta, TikTok, and YouTube that launched their respective data APIs at the event in a similar session, and this year we’re launching access to 8 different platforms within Arbiter. Earlier in the month we got a letter from a U.S. state committing to deploy Arbiter statewide for electoral safety. The same happened in an Indian state striking up a partnership for combating child trafficking including law enforcement disrupting offline networks of traffickers. India’s largest newsroom cold-DM’ed us on LinkedIn asking if we’d like to work together. I did expect it to land but have been blown away by the response. I am hard at work now on executing towards the goals, serving our partners and customers in multiple languages, across time zones, all the while trying to be the best in the world at solving this problem because that’s honestly what makes me happy. And I now have a team of 8 full-time members and my cofounder helping out along the way!
As childhood dreams go, I think we’re on the way to achieving mine.
Drop me a note telling me what your childhood dreams were and how you plan on achieving those. I promise to respond to every genuine email I get12.
Many of whom successfully got into CERN later, so there’s some correlation indicating that they do find what I write helpful.
And yes this is the first time I used that word chutzpah and actually felt like it fit the sentence excellently. I am extremely proud of myself for this.
I recently learned that’s called being a philomath thanks to memo.tv who is one of the folks involved in this tech residency program I’m currently pursuing alongside some amazing creatives and technologists in Berkeley, CA. And also this may be some signs of ADHD, idk. I never got tested and don’t have the time to, anymore.
My first and second Ph.D. advisors had their careers take a (extremely positive) turn during my Ph.D. so they ended up leaving the institute—one for an academic leadership role at another excellent dept., and the other had to leave to grow their acqui-hired startup. So I was on my third and fourth advisors at this time — due in huge part to the help from the first two advisors and a consistent co-advisor who stayed throughout these four advisors and helped develop my early work. Being a product of 5 academic parents is understandably far from a trauma-free experience, and I did nearly flunk out of my Ph.D. because the turmoil and changes came in the middle of COVID where the world was already screwed up enough. But then I see this trauma as the main reason I learnt to (1) quickly develop a research taste leading to original research ideas and (2) create independent research streams that I don’t need much external support to pursue (3) live life on the edge of losing funding and support and being potentially kicked out of the country if someone higher up decided I wasn’t doing good research and should not be funded to remain in the Ph.D. program anymore. Resilience is a skill best learned through trauma, maybe? Or maybe that’s Stockholm Syndrome talking.
Which I pursued kinda against the advice of my then-advisor, and that’s the one paper I did publish early in my Ph.D. at a decent-ish venue. I ended up taking 12 credits in a 9-credit semester and figuratively died. But I proved to my advisor that technically, at a great cost, I could do what he mentioned wouldn’t be a good idea. Unfortunately he had left the institute by then. But I really appreciate that he could have left me in the lurch and instead fought to keep me on despite my lack of academic productivity (the Adobe paper wasn’t published then so I didn’t have anything to show for all the dying) and find me other advisors. And I’m grateful to the next advisor for finding a way to take me on even amidst COVID-driven funding dry-ups.
Of course credit goes to my incredible team of coauthors including mentors, and senior leaders who helped define, enable, and accelerate my work, even developing independent pieces to help me out. You cannot understate the role of good feedback and great collaborators in driving impactful research. If you don’t enjoy working with your collaborators, your project is almost certainly doomed to fail.
I’m not saying this is a bad thing—like of course it’s great to identify online harms but there’s just no incentive for a platform to do so.
I refuse to call it X whenever I remember to refuse to call it X, out of principle and fondness.
I went to at least 6 different teams and got access to most of their datasets. But it’s never just about accessing the actual data, if you really want to build good ML models you need to appreciate the process the teams followed in acquisition, processing, and storage of the data. So I looked into the provenance, data descriptors, and use cases activated by the data. I looked at copies of live data, deprecated data, deleted data that we only kept in production, data and labels that we were planning to collect, and data that we thought of collecting but did not pursue. I examined the product docs around pipelines through which the data passed, read about current and deprecated classifiers, deployment challenges, cost optimizations, and the “bandages” different teams “stuck” on the almighty algorithm that’s now been open-sourced.
Which I am told is 2+ hrs. in ‘CEO time’. My friends were pretty surprised it was taking so long and started taking photos of us talking. I thought one of two things was true: either this guy had no idea what the state of affairs was re: data and labeling and that’s why he was asking me these questions about “where are you getting your data?” or he was a fucking genius and knew exactly what I would suffer for the 11 weeks after this chat, and was preparing me to take on the data exploration relentlessly, without giving up. You can figure out for yourself which of the two was true.
I notably exclude the hackathon-hopping, energy-drink chugging, hacker house-infestation driven tech bro ideation of “what problems can we solve to become billionaires as fast as possible”. Yes, those are an edge case where my theory doesn’t apply.
unless it’s a spammy marketer or a student asking about an internship at SimPPL. In the latter case just register your interest at https://www.simppl.org/careers and we will be in touch if there’s a fit. Follow our LinkedIn since announcements are made there first.




I've "known" Swapneel since school. But I'm not at all surprised by your life's trajectory :)
Even though we haven't spoken in years it's great to read what you write!
Great writing! The early NYU, nearly-run-out-of-funding days were tough... in retrospect, it's also a unique simulated "start-up experience" :-)