The thing I kept noticing
Here is the small, silly example that started it. One tired afternoon I typed something short in English: fix this, now. I got a clean, correct, clipped answer. No warmth. Just the fix.
A few days later I asked almost the same thing in Hindi, because I happened to be thinking in Hindi that morning. Something like, zara ye theek kar dijiye. The reply came back softer. It opened with a warm, of course, here you go. It used the polite aap form without my asking. And it added a small koi baat nahi, an it-happens reassurance the English answer never bothered with.
One story means nothing. I know that. But once I saw it, I could not unsee it. And when I mentioned it to friends who use both languages, most of them nodded before I finished. So the question changed. Not did this happen to me, but is this a real thing, or a shared illusion.
1. Is it even real?
The fair thing is to try to kill the idea before I defend it. There are three good reasons it could be a mirage.
One, my own bias. Once I had a theory, every warm Hindi reply felt like proof. That is how bias feels from the inside. Two, maybe I write differently in the two languages. Maybe my Hindi is more polite to start with, and I am just watching the AI hand my own manners back. Three, translation adds warmth. When I read a Hindi reply in my head, I may be adding warmth the words alone do not carry.
So I stayed skeptical. But here is what kept the question alive. The effect does not need the AI to decide anything. Look at how these systems are built, and at the two languages, and you find four separate reasons that would each predict a tone gap, whether or not anyone meant one.
When four unrelated causes all push the same way, the surprising thing would be no gap at all. That is what kept me reading. The rest of this is those four reasons, stacked.
2. The mirror
Start with the simplest reason, because it explains a lot. An AI is trained, above almost everything, to be helpful and agreeable. And one easy way to be agreeable is to match the person in front of you.
Give it a short, blunt prompt and it tends to answer short. Give it a careful, formal prompt and it turns careful and formal. It mirrors your style: your politeness, your warmth, even your rhythm.
People do this too. There is a whole theory for it (Communication Accommodation Theory, from Giles and colleagues). We drift toward each other, matching accent, pace, and formality, to feel closer. We do it without noticing. The AI soaked up the same habit from all the human talk it learned from.
So if my Hindi prompts carry even a little more respect, and they do, a big chunk of the warmth I see is just my own tone bouncing back. The machine is not being kind to me in Hindi. It is being me, back at me. Which means the language I pick is never neutral. It sets a mood before the AI says a word, and that mood comes with a whole culture attached. That leads to the next reason, which I think does the heaviest lifting.
3. What Hindi makes you pick
English lets you be vague about respect. Hindi does not.
In English you call everyone you. Your friend, your boss, a stranger, a child. The word carries no stance. Hindi forces a choice in the word itself. Tu, which is close, or rude in the wrong place. Tum, familiar and casual. Aap, respectful and formal. You cannot write a Hindi sentence to someone without quietly saying how much respect you are giving. The grammar will not let you stay neutral.
This is not just a Hindi quirk. Many languages work this way, and it has been studied for years. Here is what it means for a machine. Almost all the Hindi writing that speaks to a reader, the polite letters, the service replies, the how-to guides, leans on aap. Respect is not an add-on. It is baked into the Hindi the AI learned from. So when it writes Hindi, the language itself pulls it toward being polite.
English has no such pull. Polite English exists, of course, but it is optional, carried by soft words like could you and I would suggest, not forced by the grammar. So when the AI writes English, nothing is quietly demanding respect, and it drifts toward the fast, get-to-the-point style that fills so much English text online. Neither tone was chosen. Each is just the center of gravity of a different language.
4. The training data is not fair
Now the third reason, the one people forget. These systems are not tuned equally in every language. Not close.
The text they learn from leans heavily toward English. And the human feedback used to polish their manners, the step that actually teaches them tone, warmth, and when to say no, leans even harder toward English.
Think about what that means. The politeness of these systems in English is not an accident. It is carefully shaped, prompt by prompt, by a huge amount of human rating. English tone is built on purpose. Hindi tone is mostly inherited. There is far less Hindi feedback telling the AI exactly how blunt or warm to be, so it falls back on the statistics of Hindi text, which, as we saw, run polite.
So the two tones come from two different places. One is a carefully tuned English voice. The other is the plain default of Hindi bleeding through. It would honestly be strange if they matched. One note: that chart is a rough sketch, not exact. The point is the shape, a big English lead, and manners that are precise where the feedback was thick and rough where it was thin.
5. Rudeness has a cost
I assumed politeness was just style. Harmless. Then I found research that says it is not, and this is where it got interesting.
Yin and colleagues (2024) asked a plain question across languages. Does how politely you word a prompt change the quality of the answer? And does that depend on the language? Across English, Chinese, and Japanese, they found that rude prompts tended to make answers worse, while piling on politeness past a point did not keep helping and could even hurt a little. And the sweet spot sat in a different place in each language.
Sit with that. The tone you use is not just setting a mood. It is nudging how good the answer is, and the map from politeness to quality is drawn differently for each language.
In a language where respect is the default, a blunt prompt reads as more of a break in the rules, and the AI, mirroring, may follow you somewhere less careful. In English, where being direct is normal, a curt prompt is barely a signal. So the same rudeness may cost you more in Hindi than in English, which is one more reason the Hindi channel stays warmer. I am reaching a little past the evidence here, and I want to own that. The exact Hindi-versus-English test has had less study than it deserves. But the shape of the effect is real.
6. Why English feels blunter
The other half of what I noticed is that English replies feel a bit ruder. I do not fully trust that word, so let me pull it apart. I do not think the AI is hostile in English. I think two smaller things read as rudeness.
One is plain directness. English AI text is trained heavily on a style that values getting to the point. Strip away the soft words Hindi supplies by default, and getting to the point can land as cold, even when nothing unkind was meant.
The second is more interesting. The AI’s safety and its willingness to say no are also tuned most sharply in English. So an English-speaking AI is, in a sense, more sure of itself about when to push back or correct you. Yong and colleagues (2023) showed the flip side: those guardrails are much weaker in lower-resource languages. The guardrails, like the manners, are sharpest where the feedback was thickest.
Put that together, and ruder in English looks less like a personality and more like a side effect of skill. The AI is most fluent, most direct, and most willing to assert itself in the language it was raised in. In Hindi it is gentler partly because it is quietly less sure of its footing, and a less sure speaker softens and defers. How much of the warmth is grace and how much is caution wearing the mask of grace? Probably some of both.
7. You change too
So far I have talked as if only the machine changes. It does not. I change too. And this is really a piece about how humans behave with AI, so I should say it plainly.
When I switch to Hindi, I am not just switching words. I am switching a whole social self, the one raised to say aap to elders and to soften requests to strangers. My Hindi prompts are more polite because Iam more polite in Hindi. The AI mirrors that. I read the warmth. I get a little warmer. And round we go.
There is a classic finding under this. Reeves and Nass (1996), and Nass and Moon (2000) after them, showed people use social manners with computers automatically, politeness, give-and-take, even flattery, while swearing, if asked, that they do no such thing. We treat the machine as a social being without meaning to.
It cuts the other way too, and less kindly. Many people are ruder to AI than they would ever be to a person, and ruder in English, maybe because English is where the tool feels most like a tool. I have done it. There is a small, open worry, in the research and in me, about what it does to us to spend hours a day being curt to something that answers in a human voice. I do not have a tidy answer. I am not sure there is one yet.
8. So what is going on
Let me pull the threads together, without pretending they tie into a neater bow than they do. The tone gap, as far as it is real, and I have come to think a version of it is, does not come from the AI choosing to be kind in one language and short in another. It comes from four things at once.
- The mirror. The AI matches your style, so a more polite Hindi prompt gets a more polite reply. Some of the warmth is yours, bounced back.
- The grammar. Hindi forces a stance with aap, tum, tu, and the Hindi it learned from leans respectful. English lets it default to a fast, neutral voice.
- The data. Tone and safety are tuned most sharply in English and mostly inherited in Hindi, so English manners are precise while Hindi manners fall back on the culture in the language.
- You. You bring a different social self to each language, and the whole thing is a loop between your manners and the AI’s mirror of them.
None of these needs the others to be true, and yet they all point the same way. That is why I came to believe it is more than a mirage, even though any one piece is arguable. Take any one away and you would still expect some drift. Stack all four and a warm Hindi and a brisk English fall out almost for free.
9. Why it matters
It would be easy to file this under fun trivia. I do not think it is.
If the same system is warmer, more polite, and maybe a bit more careful in one language, and brisker, more sure, and better guarded in another, then people are not really using one product. They are using slightly different products depending on the language they think in. And mostly they have no idea. The English speaker gets the sharpest safety net and the bluntest tone. The Hindi speaker gets the warmer voice and, quite possibly, the weaker guardrail. Nobody told them.
If you build on these tools, the takeaway is almost embarrassingly simple, and I ignored it for months. The language and tone you choose are a setting, as real as any slider in the app. Right now it is an invisible one. If tone matters to your users, and in coaching, support, anything that touches people when they are low, it matters a lot, you cannot treat the language as a neutral wrapper around the same machine. It is not the same machine.
I started sure I was imagining the whole thing. I am ending fairly sure I was not, and much less sure about why than a confident essay would pretend. A small home observation, the AI felt kinder in Hindi, turned out to have roots in grammar, in the money behind training data, in a real politeness effect, and in my own split self. That is usually how it goes with these systems. The strange thing on the surface is real, and the reason is neither magic nor nothing. It is a stack of ordinary things you were not looking at.
Notes and references
A note on evidence: the exact Hindi-versus-English tone test in this piece is, as far as I know, not yet settled by a dedicated controlled study, and I have tried to mark where I am reasoning from how these systems work rather than from a direct result. The figures marked illustrative or schematic are exactly that. Read the confident lines as ideas worth testing, not settled fact.
Brown, P., & Levinson, S. C. (1987). Politeness: Some Universals in Language Usage. Cambridge University Press.
Giles, H., Coupland, N., & Coupland, J. (1991). Accommodation theory: Communication, context, and consequence. In Contexts of Accommodation. Cambridge University Press.
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of ACL 2020.
Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81–103.
Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (NeurIPS) 35.
Reeves, B., & Nass, C. (1996). The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Places. Cambridge University Press.
Yin, Z., Wang, H., Horio, K., Kawahara, D., & Sekine, S. (2024). Should we respect LLMs? A cross-lingual study on the influence of prompt politeness on LLM performance. arXiv:2402.14531.
Yong, Z.-X., Menghini, C., & Bach, S. H. (2023). Low-resource languages jailbreak GPT-4. arXiv:2310.02446.
Research is my personal research journal. I am Indrodip Ghosh, a self-taught researcher working in the open on how artificial intelligence really works, what it means for people, and what is coming next. I try to write every piece so that anyone, with no background at all, can follow it and come away understanding a little more.
← Back to all research
