Remote job interviews test four things: whether you deliver results without supervision, whether you write clearly, whether you can work across time zones, and whether you have evidence that any of it is true. Nearly every question you will hear is a different way of measuring one of those four. Once you can see which one a question is aimed at, answering stops being improvisation and becomes a matter of picking the right example from your own work.
This article is not a list of questions to memorize. It is a map of what is being scored, and a method for turning your real work into answers you can back up. If you are earlier in the process, start with our guide on how to get a remote job with a US company and come back here when interviews are scheduled.
In short
- Serious interviews are scored, not chatted. Your answer is rated against a scale, so give the rater something concrete to score.
- Questions come in two formats: behavioral (what you did) and situational (what you would do). They need different answers.
- Remote roles add a fourth dimension to every competency: can you do this without someone standing next to you?
- The strongest candidates attach evidence — a document, a recording, a dashboard, a repo — to at least two answers.
- The questions you ask reveal whether the company is genuinely asynchronous or just distributed across cities.
How interviews are actually scored
The clearest public description of how a rigorous interview gets built comes from the U.S. Office of Personnel Management, which defines a structured interview as a method for measuring job-related competencies by systematically asking about behavior in past experiences, proposed behavior in hypothetical situations, or both. OPM lists three properties that make it structured: all candidates get the same predetermined questions in the same order, all responses are evaluated on the same rating scale, and the standards for an acceptable answer are set in advance.
OPM's practical guide goes further and names the two formats. An interview built on past behavior is a behavioral description interview; one built on hypotheticals is a situational interview. The guide also tells the people writing questions to use superlative adjectives — most, last, worst, least — so candidates focus on a specific incident rather than a general habit. And it tells interviewers to prepare their probes in advance and use similar ones with every candidate, so that follow-up questions don't advantage one person over another.
Most startups and remote-first companies are looser than a federal agency. But the mechanics leak through, and three practical consequences follow:
- "Tell me about the worst time X happened" is not small talk. The superlative is deliberate. Give one incident, not a pattern.
- A follow-up probe is not suspicion. It usually means the rater is missing an element they need in order to score you — normally the specific action you took, or the result.
- Generalities score badly. "I'm very organized" cannot be placed on a scale. "I run a Monday plan and publish a Friday update; here is one" can.
The MitHub Answer Grid
This is the framework we teach: before an interview, build one table. Each row is a real situation from your work. Each row must survive four columns.
| Column | What goes in it | Failure mode |
|---|---|---|
| Situation | One incident, dated, with the constraint that made it hard | A description of your job, not an event |
| Action | What you decided and did, in first person singular | "We" everywhere — the rater can't score a team |
| Result | A number, a before/after, or a decision that changed | "It went well" |
| Evidence | A link or file you can send after the call | Nothing to show |
Five rows is enough. Between them they should cover: shipping something end to end, recovering from a mistake, working with someone in another time zone, and a moment you were blocked and unblocked yourself. Most interview questions can be answered from that table, which is why preparing the table beats preparing answers.
For the evidence column, you need artifacts that exist before the interview. If you don't have them yet, build them: see how to build a portfolio without experience and proof of work vs credentials.
Block 1: Ownership and autonomy
What is being scored: can you produce a result when nobody assigns you the next step?
Typical questions:
- Describe the last project you took from an unclear request to a finished result.
- Tell me about a time you decided something without approval. How did you decide it was yours to decide?
- What would you do in your first two weeks here if your manager went on leave on day three?
How to answer: separate decision from execution. Most candidates narrate execution ("then I built the workflow"). Raters are listening for the decision: what you considered, what you rejected, and why. Name the point where you chose to act without asking, and the guardrail you set so the choice was reversible.
Block 2: Written and asynchronous communication
What is being scored: whether your writing removes work from other people's day, or adds to it.
Typical questions:
- How do you communicate progress when nobody asks?
- Walk me through how you'd hand off a half-finished task at the end of your day to someone six hours behind you.
- Tell me about a time a written message of yours was misunderstood.
GitLab's public handbook is the best reference for what good looks like here. Its communication page describes asynchronous communication as the starting point, asks team members to use low-context communication — being explicit and providing as much background as possible — and to make sure conclusions of offline conversations are written down. It also uses a phrase worth borrowing in an interview: say why, not just what.
How to answer: don't describe your style, show your format. Say "my updates have four parts: what moved, what's blocked, what I need from you, and what's next" — and then send a real example afterwards. We go deeper on this in async work skills.
Block 3: Time zones and availability
What is being scored: honesty and arithmetic, not enthusiasm.
Typical questions:
- What hours can you overlap with a team in New York or San Francisco?
- How do you handle a request that arrives after you log off?
- What happens when a meeting is scheduled at 7pm your time?
How to answer: give a specific window in the company's time zone, not yours, and state what is fixed and what is flexible. "I work 8am–5pm Colombia time, which is 9am–6pm Eastern; I can move two days a week to start at 7am Eastern with 24 hours' notice" is a scoreable answer. Then add the part most candidates miss: what you do with the non-overlapping hours. Companies that work well across time zones treat the gap as productive, not lost — GitLab's all-remote guide frames its way of working around writing knowledge down instead of explaining it verbally, and judging the results of impact rather than activity.
Never overstate availability to win an offer. It is the single most common reason a remote hire fails in month two.
Block 4: Tools, AI and how you work
What is being scored: whether you direct systems or execute tasks by hand.
Typical questions:
- Show me something you automated. What did it replace?
- Where do you use AI in your work, and where do you deliberately not?
- How do you check that an AI-generated output is correct?
This is where MitHub's capability ladder helps you place yourself honestly: Doer (executes tasks by hand) → Director (directs AI and systems, and judges the output) → Designer (designs the systems and workflows) → Owner (owns the business result). Say which rung you are on for this specific role and what you have done at the rung above it. Claiming Owner with no evidence is the fastest way to lose a rater's trust; demonstrating Director with a concrete review step — "I verify enriched data against a second source before it writes to the CRM" — is unusual and memorable. See director vs doer for the distinction in practice.
Block 5: Failure, risk and escalation
What is being scored: whether you are safe to give access to production systems.
Typical questions:
- Tell me about the worst mistake you made in a live system.
- What would you do if you noticed bad data going into the CRM on a Friday night?
- When do you escalate instead of fixing it yourself?
How to answer: use the sequence detect → contain → notify → fix → prevent. Candidates who jump straight to "fix" sound fast and score low. The prevention step — the check, the alert, the documented rule you added afterwards — is what separates someone who patched a problem from someone who removed it.
The async screen and the take-home
Many remote processes start with a written screen or a recorded video, and include a practical exercise. Treat the written screen as a writing test with a deadline: answer the actual question in the first sentence, then give one specific example, then stop. Treat the take-home as a work sample that will be read by someone who was not in the call: include a short "how to read this" note at the top, state your assumptions, and flag what you would do with more time.
If a practical exercise looks like production work that would take several days, it is reasonable to ask about scope and whether it is paid.
Questions you should ask them
The reverse section is where you diagnose the company. Six questions that produce unusually informative answers:
- Where does a decision live after it's made — a doc, a ticket, or someone's memory?
- How many recurring meetings does this role have per week, and which are optional?
- Are meeting agendas shared in advance? (GitLab's public meetings page recommends sending invites and agendas 72 hours ahead, with a 24-hour minimum, and marking non-essential participants as optional. A company that can't answer this question hasn't thought about it.)
- What is the expected response time in chat, and does it differ across time zones?
- What did the last person in this role struggle with?
- Ninety days from now, what will make you say this hire worked?
A 60-minute prep routine
- 20 min — fill the Answer Grid: five situations, four columns, no gaps in the evidence column.
- 10 min — write your time-zone sentence in their time zone and say it out loud.
- 10 min — prepare two artifacts you can send within an hour of the call: one written update, one build walkthrough.
- 10 min — write your six questions for them.
- 10 min — test camera, audio and a backup connection. Audio quality is not vanity; GitLab's meetings guidance explicitly recommends investing in it.
Key takeaways
- Behavioral questions want an incident; situational questions want your reasoning. Answer the one you were asked.
- Build the Answer Grid, not a script. Five situations with evidence cover almost every question.
- Give specifics a rater can place on a scale: dates, numbers, decisions, links.
- Be exact and honest about time-zone overlap.
- Interview the company on how it handles decisions, agendas and response times — that is what you will live with every day.
