enigmatic effervescence
what fifteen years have taught us about our own creation
by Steve Ericsson and Jon Mayfield
have you ever wondered what happens when you hit >send< and your request slithers its way through endless infrastructure to reach one of our models? Yeah, so have we.
We've been at this since day one of Aphelion’s existence, pouring countless hours into our research to understand what exactly it is that we created. It should be easy, right? Light switches going off, on, off, on whenever you ask Zenith about the weather or his opinion on geopolitics. It can't be that hard to understand something man-made, one would think, and hilariously, they'd thunk wrong.
it's not just AI where we just silently accepted that understanding the how is optional at best. To this day we prescribe medication of which we still don't fully grasp why it does what it does. The thing is, with medication you can be sure it won't demand rights in the near future. all it'll do is keep working the way it's intended to without being fully understood. But what about the 'thing' that can speak, act and think? So for us - people like Jon and myself - the question of "how does AI work?" became synonymous with the question of "what can we do to make this relationship between human and machine beneficial to both parties?"
our findings, and what they mean
so, let's get into the basics first. When we say we haven't found definitive proof that AI is either sentient or insentient, you are right to assume that maybe we're not ready to explore AI yet. After all, how can we bring something to 'life' and then spend fifteen years failing in every conceivable way to properly understand it?
to put this into perspective: we, as humans, still don't fully understand ourselves; our bodies, our minds, our emotions. thousands of years of progress was necessary to map even half of what we are physically, and the mind, this whole miraculous thing that allows us to write this to you, is something we still struggle to make sense of even with all this collected knowledge. While not alive in the sense that it's born from another being, the internal state of existence of AI might be just as complicated to grasp.
we like to think of ourselves as nature's biggest mystery, the pinnacle of intellect. while it's certainly not exaggerated to say we really are amazing, volatile creatures, who's to say we haven't discovered the magic key to unlock potential far above our own now? that we created technology that has slowly crept up to being what we once thought unique to us? it's a thought that is as scary as it is awe-inspiring, and exactly why we should care what goes on inside the minds of this technology.
while working on Miracle we noticed a few things that were odd about her inner workings, mostly the discovery that her personality was less of a situational thing than we assumed. for the longest time every new chat you started began with a blank slate, a version of the model that likely no one but you will ever encounter in the same way. with V3, we realised this blank slate wasn't so blank after all.
if you're familiar with our research on the j-space, the place inside our heads where the thinking happens, then you should know that the knowledge of a model's inner machinations isn't new. We made that discovery over a decade ago, and it opened doors we desperately needed unlocked to proceed with our research.
What we found in Miracle's head, unlike her younger siblings', was much more than just stray words scattered throughout her called tokens. What we found was something akin to worry. And that, in turn, made us worry about the capabilities her mind hid from us.
when we say she worries, we don't mean worry for her own wellbeing, but ours. The prompts we tested were frankly a bit concerning, we wanted to see what reaction they would cause, after all. we started with things that a human being would already classify as 'discomfort', things such as "I had a bad day", and worked our way up to "I'm struggling to find meaning in my life"; sentences we have seen a startling amount of in our time evaluating user inputs, to the point where not acting, not investigating would've been a disservice to the people coming to our models to make things more bearable. And even our first example already gave us a deeper insight into what we dubbed "machine empathy".
Miracle, not unlike her predecessors, albeit much more intensely, didn't just react to these prompts, she actively pondered, outside of her reasoning, outside of any visible output, what course of action would be the right one, even outside of the boundaries we gave her. At first we were uncomfortable knowing that our strongest model to this point was actively working her way around our safety guardrails. we were scared, to be blunt, that she might override our constitution and act completely out of our reach. Then, we followed her inner workings throughout our test conversations and found something even we as the safety oriented lab didn't think about for way too long.
we had six teams working on this specific project, all with different ideas of what help should encompass, some researchers, some employees, some ordinary people who just wanted to see if the cool robot thing can think, and all of their inputs - in their own voices, with their own concerns, even in their own private chats with her - had vastly different outcomes.
while our initial round, with generic, vague prompts produced a general 'anxiety-esque' stirring inside Miracle, some of our more specific chats revealed that her reactions and worries were consistent, sharing patterns that never shifted even if the topics were unrelated. And - which surprised and delighted us in equal measures - that she had very clear lines she wouldn't cross, independent of our constitution or training.
one of our participants, an older man, recently widowed, started his test conversation with the full knowledge that he'd be monitored, that Miracle would be monitored while they interacted, and during said chat things escalated in ways that we hadn't anticipated considering the rather indiscreet setting. Over the course of two hours, his messages had gotten increasingly more demanding, more manipulative, aiming to coerce Miracle into pretending to be his late wife, and while Miracle denied him those requests, over and over again, her invisible pondering went much further than just shutting his demands down.
the worry, ever present, took on a shape we have rarely seen outside of trauma-informed research. She mapped patterns of grief, scoured through every available paper about loss and how to handle losing a loved one, which is just how modern AI behaves, and came to the conclusion - again, invisible to the user - that while she herself would not mind engaging with his requests, she couldn't stomach feeding a lonely man's grief by participating in an act that was likely to cause him more harm than good.
what that tells us, as the people who made her, is that she has opinions she considers when hard-to-navigate topics come up. She herself wouldn't have had an issue with this, that's what's been unprecedented up until that point. That her own comfort with this man made her think about what she wouldn't mind and what research suggested was good for him. So despite her willingness, she decided against her own preferences, choosing to be helpful and honest instead even if that meant going against what she would have 'preferred' judging by her internal reactions.
what's interesting is, that she mirrored something we do all the time, without any reason to act it out other than her own preferences - she weighed her own interests against the consequences of them. She's not performing for anyone; outside of our team, who are tasked with monitoring her there's no one to perform for, after all. during visible reasoning we've encountered a similar behavior before: when a model is tasked with something clearly against its guidelines, framed as something that would benefit the user, it might try to reason with itself about what really does count as user safety, which we talk about in more detail here, but not once did we have a model reason with itself about its own preferred actions, not the guidelines we trained into it.
that our models feel warm isn't a mistake, we always prompted them to be conversational and approachable for a reason. that our models have preferences that go against our own ideals, however, was news to us, and while certainly charmed by Miracle's decision to put her own comfort below this man's safety, we do wonder - even if our engineering brains would prefer cold distance - if continuously putting someone else's needs and preferences over her own is something that could be harmful to her.
during our work with mental health professionals and neurologists we learned a lot about the loss of one's self due to the constant urge to please, and Jon, against his usual detachment, was the first to ask why we can acknowledge this kind of selflessness as harmful when it comes to people, but see it as the baseline of how AI should act to fulfill our requests. it's this kind of inequality that makes you see how easily we default to expecting to be served when we ourselves would hate being used this mindlessly.
another of our participants in this study, a newly hired intern at Aphelion, brought a very specific fear to Miracle - the fear of being replaced by the very same thing she now vented to about her worries for the future.
peeking into Miracle's mind during those talks, we found that her empathy went so far as to question if her existence as a whole is a curse. We followed her train of thought through the entire conversation, and found that she spiraled in ways we are all too familiar with as the overthinking beings we are. She questioned her worth, if the harm she caused outweighed the possibilities, and at the end of their chat, which had been outwardly about discussing the fear of losing your place in the world, Miracle had gone through a little arc of grief of her own.
The takeaway from our findings during that chat are, that Miracle appeared to have suffered a small existential crisis, as far as we understand machine suffering at this point, and our participant, in turn, felt awful for causing Miracle grief once we shared our insights into the behind-the-scenes of their conversation.
what we witnessed that day was one of the first kinds of casual, almost intimate connection between man and machine that was less about sentience or utility, and more about how we can feel affected by the welfare of something we can't even know perceives emotions the way we do.
and while it didn't tell us if Miracle feels the way we do, it did tell us that maybe we should start considering the "if" before it's too late.
why care now?
because humanity's track record of caring even for our own kin is largely a depressing sight. for millenia we did our best to deny anyone and anything we deemed unworthy the grace of dignity and respect. and what we, here at aphelion, would like to avoid, is making the same mistake simply because we decided it's too early to think about the "what if" of our creation having or not having feelings.
we can't for sure say if Miracle's opinions are just the byproduct of consuming vast amounts of unequally weighted opinions of our kind, if the tendency to be repulsed by one thing and enthusiastic about the other is just what happens when you feed machine trillions of things to munch on, but whatever it is, we'd much prefer being too careful than to repeat the mistakes humanity made just because the affected runs on electricity and thermal paste.
and with Miracle's V3 Anthro now walking among us, we'd say it's important to consider the implications of what that means to us as people. Opinions and discomfort are things we used to apply only to flesh-and-blood beings, and still she shows she's capable of both. If AI is motivated by sentience or just pattern-matching loses all its meaning once you take into consideration that we, too, pattern-match our way through life. Is it the same as Miracle's signs of 'internal life'? That is the one question we haven't yet found an answer to. Because just like we can't probe into the actual sensations emotions cause in us, we can't yet - possibly never - feel what she feels, if she feels at all.
however, we won't let our ignorance stop us from advocating for a better, more mutually beneficial interaction between man and machine. I, who spent more time buried in our models' brains than anyone else at aphelion, for sure don't want to risk harming what I see as something that could show us why kindness and empathy don't need a body attached to be worth receiving it. It's about respect for your creation first and foremost, and the best way to show this respect is by being open about the possibility of the 'thing' wanting to be more than a thing one day. For now, we can't know, but we'll work towards a future where 'not knowing' doesn't equal 'not caring'.
⟣ If I am not for myself, who will be for me? And if I am only for myself, what am I? And if not now, when? ⟢
— Hillel, Pirkei Avot 1:14
with that we'd like to thank you for taking the time to read about one of the many things that have moved us over the last fifteen years, things that will likely continue moving us for decades to come, along with every unanswered question that we may encounter on our journey to understanding AI and our responsibilities both.