macOS screen recorder
Framedly
A screen recorder that re-shoots the take.
It records the screen and the pointer as separate tracks, then renders the video from both — pushing the camera in on every click, smoothing the cursor, and setting the result padded, rounded and shadowed on a generated background.
Coming soon Requires macOS 15 or later
Built with ScreenCaptureKit, Core Image and AVFoundation. No dependencies.
- macOS 15+
- H.264 · HEVC · ProRes
- Up to 4K
- Animated GIF
- On-device captions
- No account
- No server
01The idea
A plain recording is a wide shot that never moves
What makes a recording look directed is camera work: pushing in when something is clicked, holding, easing back out — and a cursor that glides instead of twitching. Doing that afterwards means knowing where the pointer was and when it was clicked, which a video file no longer contains. So Framedly records the pointer as its own data track and hides the real cursor. Everything visible in the final video — the cursor, the zooms, the framing — is drawn at render time from that data, which is why a zoom stays sharp and the cursor keeps one size however far the camera pushes in.
02How it fits together
Capture, solve, render
A recording is a directory holding screen.mov, camera.mov, the audio files, the captured cursor bitmaps and project.json. The media is plain files, so a recording survives even if the document fails to load, and a project written by an older build opens with the newer settings at their defaults rather than failing.
Once clips can be trimmed, cut and re-timed, "3.2 seconds in" means two different instants depending on which timeline you mean. Every recorded track is stamped in source seconds and the compositor is handed edited seconds; TimeMap is the single place that conversion happens. Because the solver iterates over edited time, a clip played at 2× moves the camera at the same physical rate as one at 1× rather than twitching twice as fast.
Preview and export both run the same compositor — one as an AVPlayerItem, one through an asset reader — so what you scrub through is produced by the code that produces the file. There is no second render path that can drift out of agreement.
-
Capture
What the machine can only know while it happens
- screen.mov
- camera.mov
- pointer + clicks
- keystrokes
- audio + loudness
-
Solve
The whole recording is known, so the answer can look ahead
- TimeMap — source ↔ edited
- MotionSolver
- FaceTracker
- MotionPath
-
Render
One compositor, so the preview cannot disagree with the file
- FrameCompositor
- AVPlayer — preview
- AssetWriter — file
- ImageIO — GIF
03Automatic zoom
The camera anticipates, it doesn't chase
Because the whole recording is known up front, the camera does not have to be causal. It ramps in from wherever it was to the click point on the same eased curve as the zoom, so both finish together and the clicked element is centred exactly as the click lands — and the ramp starts about 0.55 s before the click, so the camera is already moving when it happens. A damped follower could only ever arrive late.
Then it holds, handing off to a lazy follower with a deadzone so the shot tracks drags and later clicks without drifting on every pointer twitch. Its time constant tightens as the pointer gets further ahead, with a floor so catch-up never becomes a whip-pan, and a hard containment box as a backstop. Finally it ramps out, recentring so the shot lands neutrally at 1×.
Cursor smoothing runs two cascaded one-pole filters — stable at any timestep, with the second pole removing the velocity discontinuity a single pole leaves. Smoothing introduces lag, which would land clicks next to their target, so the pointer is snapped back onto its true position in a short window around every click.
04Motion blur
Integrated, not smeared on
A camera that moves without blurring strobes, so the shutter is modelled rather than approximated: the frame is the source resampled at a dozen instants across the time one exposure is open, and averaged. The camera's move is an affine, so that average is the exposure — there is no filter radius to calibrate against the frame rate, and nothing to get wrong where the picture meets its own edge.
A directional blur applied afterwards cannot express what a zoom does. A zoom's motion is radial: content beside the point being zoomed into barely moves while the far corner travels tens of pixels, so one angle and one radius describe neither. Integrating gets that for free, and it is what keeps the cursor honest — the pointer is sampled at the same instants as the picture, so its smear and the screen's agree instead of being two effects tuned to match.
Taps are spaced a fixed fraction of the frame apart rather than a fixed number of pixels, so a held shot costs one array lookup and no drawing at all, and a 4K export does not come out sharper than the preview it was judged in.
05Camera and privacy
The face mask is solved, not detected live
Hiding your face in the webcam bubble looks like a job for a per-frame detector, and it is not: one running under the compositor would cost every frame twice, and it would fail in the two ways that make a privacy mask worthless — it shivers frame to frame, and it drops out the instant you turn your head, uncovering for a few frames the face it exists to hide.
So Vision runs over the camera file once, at 10 Hz, and the result is solved into a fixed-rate track. A dropout has a known-good box on both sides, so it is filled from both, and the box swells through the middle of a gap by more the longer the gap ran — a blink costs nothing, while a two-second absence is covered generously rather than precisely wrongly. Smoothing runs one pole forward and the same pole backward, so the phase the first pass adds the second removes, and a deadzone pins the box outright while your head is still.
If detection finds nobody, the mask covers the whole overlay rather than doing nothing. A privacy control that silently no-ops fails invisibly in the preview you are checking and permanently in the file you publish.
06What it records
Screens, windows, an area, or your phone
-
Any source
Whole screens, a single window, a drag-selected area, or a connected iPhone or iPad — with the model and colour detected and framed.
-
Webcam on the same clock
Written to its own file through a data output rather than a movie output, because that stamps frames on the host clock ScreenCaptureKit uses. The difference between the two first frames is what keeps lips in sync.
-
Audio that does not drift
A microphone takes half a second to spin up, and AVAssetWriter starts the session at whatever it is handed first. The lead-in is written out as real silence, and so is any gap left by a dropped buffer.
-
Keystrokes
Key caps that fade in and out, shortcuts-only or everything. Input Monitoring only applies from the next launch, so the app says so rather than silently recording nothing.
-
Your actual wallpaper
Read as pixels through ScreenCaptureKit rather than from a file, because wallpapers moved behind providers and the obvious API now hands back a stock picture most people have never seen.
-
Captions, on device
Transcription with word timings, so the spoken word lights up as it is said. Nothing is uploaded; the accuracy is whatever the system model manages for the language in use.
07Export
A destination is a specification, not a resolution
Picking YouTube 1080p or YouTube 4K sets the container, profile, frame shape, GOP length and entropy mode at once — including the 16:9 canvas, because a 16:10 screen exported as-is arrives as a file YouTube has to pillarbox, and by then the bars are pixels. Editing any individual control drops the picker back to Custom rather than keeping a name that no longer describes the file.
Speed re-times the whole edit and scales the composition rather than the encode, so the preview plays at the chosen rate too — a 2× export you can only check by exporting is a setting that gets shipped wrong. Past 16× the rate stops being a multiplier and becomes a mode: the cursor, click highlights, zooms, key caps, captions and recorded sound are all left out, because everything that makes a screen recording readable is defined in human time. The ceiling is 240×.
Exports do not need the editor. A .framedly is a directory of separate tracks, and folding it into one file is a pure function of what is already on disk — so the same exporter the button runs will run from a terminal, write the file beside the project, and report progress on stdout.
Framedly --export ~/Movies/Framedly/Name.framedly --preset youtube-1080Framedly --export ~/Movies/Framedly/Name.framedly --speed 2xFramedly --export ~/Movies/Framedly/Name.framedly --speed 240x
08Verifying it works
The parts that have an arithmetic consequence are asserted
Screen capture needs a human to grant permission, but everything after capture is data in, data out. The self-test synthesises a recording — camera track included — runs the real export path over it twice, and checks what can be checked: clicks land in the middle 60% of frame, the smoothed cursor lands on each click, a 2× clip halves the timeline and a 2× export rate halves it again, a cut instant has no place on the timeline at all, a vertical canvas exports 720×1280, and silence detection finds the silence that was planted for it.
Motion blur is asserted over a frame of hard stripes at the instant the camera moves fastest: it has to soften that frame, leave a held one bit-identical, and not shift the mean brightness — because averaging premultiplied pixels the wrong way still looks like blur and only gives itself away as a darkening. Sample frames are written out for the parts only an eye can judge.
build/Framedly.app/Contents/MacOS/Framedly --selftest /tmp/selftest
09Not built
What it deliberately does not do
- Shareable upload links. There is no server, so there is nothing to upload to and nothing of yours to keep.
- Multi-track compositing of several separate recordings.
- Transparent backgrounds outside ProRes 422 — it is the one export that carries an alpha channel.
- Cloud transcription. Captions are on-device, so their accuracy is whatever the system model manages for the language in use.
10Everything else
What's in the editor
- Background — 21 generated mesh-gradient wallpapers, gradients, colours, images, your actual desktop picture, or transparent; blur and grain.
- Shape — padding, rounded corners, an inset frame, directional shadow, edge highlight, and iPhone/iPad device frames with a choice of body colour.
- Canvas — Auto, Wide 16:9, Vertical 9:16, Square, Classic 4:3 or Tall 3:4, plus an interactive crop. Zooms re-frame themselves to whatever shape you pick.
- Timeline — drag trim handles, split at the playhead, per-clip speed, delete, and a waveform drawn from the same loudness envelope the silence detector reads.
- Silences — cut them out or speed through them, by threshold and duration.
- Zoom — automatic on clicks, or placed by hand; per-segment level, anchor, easing and curve.
- Cursor — size, smoothing, zoom compensation, click highlights, auto-hide when still, loop back to the start position, and always-use-the-arrow.
- Camera — circle, rounded or square, any corner, size, margin, zoom, mirror, shadow, and a face mask that blurs, pixelates or blocks out your face and follows it.
- Captions — on-device transcription with word timings, so the spoken word lights up as it is said; size, position, colour, plate and words per line.
- Keystrokes — key caps that fade in and out, shortcuts-only or everything.
- Audio — per-source levels, plus a looping music bed with fades and ducking under the voice.
- Presets and export — five built-in looks and your own; H.264/HEVC/ProRes to 4K, animated GIF, copy to clipboard, and undo/redo throughout.
Framedly
A screen recorder that re-shoots the take.
Coming soon Requires macOS 15 or later
Built with ScreenCaptureKit, Core Image and AVFoundation. No dependencies.