Start from the assumption that the candidate on your remote technical interview has a model open on a second screen. Not because most people cheat, but because designing an assessment that only works if nobody does is how you end up trusting a result you have no reason to trust. Once you accept the assumption, most of the standard remote interview stops being worth the hour.
The formats that died first
Anything with a single correct answer that can be typed is gone. Algorithm puzzles, the standard take-home, the tidy implement-this-function screen, the multiple-choice technical quiz. These were never great predictors of senior performance, but they were cheap and they weakly correlated with something. They now correlate with access to a tool that everyone has. A candidate who produces a clean, tested, well-structured solution in twelve minutes has told you nothing except that the problem was a common one.
Surveillance is the wrong instinct
The reflex is to lock things down: proctoring software, a locked browser, eye-tracking, a second camera pointed at the room. Some of that has a place in a high-stakes final stage. As a general answer it is bad, because it degrades the experience for exactly the senior candidates you cannot afford to annoy, it produces an adversarial dynamic in the first hour of a relationship, and it still does not tell you whether the person can do the job. You end up with a well-monitored recording of a test that was already measuring the wrong thing.
What still works: make the subject their own history
A model can solve a generic problem. It cannot tell you why the person you are talking to chose Postgres over the thing their team already ran, in a system that is not public, three years ago, under a deadline that no longer exists. Anchor the assessment in the candidate's specific claimed work and follow the answers wherever they go. The questions cannot be prepared in advance because they are generated from the previous answer. There is no way to look up a conversation about your own decisions.
What still works: watch the reasoning, not the artifact
Give someone a piece of code with a real defect in it and ask them to talk through how they would find it, not to fix it. Ask what they would instrument first and why. Ask what they would expect to see if their hypothesis were wrong. The valuable part is the order in which someone reaches for things and how they behave at the moment they realise they are on the wrong track. Both are extremely hard to outsource in real time, and both are close to the actual job.
What still works: make being wrong safe and then watch
Give a candidate a scenario with a genuine trade-off and no clean answer, take the opposite position, and push. You are looking for whether they can hold a view under pressure, change it when the argument is good, and tell the difference between the two. That single behaviour predicts more about a remote senior hire than any test score, and it is the one thing an assistant on a second screen cannot supply, because it is not a question of knowledge.
The uncomfortable conclusion
All of the formats that survive are live, adaptive and conducted by someone senior enough to follow an unscripted technical conversation. That is more expensive per candidate than a take-home, which is why so many teams are still running the cheap version and quietly discounting the result. You either pay for real assessment somewhere in the process, or you pay for it later, at a much worse exchange rate, after the hire.
We assume assistance and assess accordingly, which is why our sessions are live, adaptive and anchored in the engineer's own work. See how the assessment works.