The co-writer learns what you rate and what people replay, fingerprints your voice, then grades your 500 Kismet songs, keeps the good parts and writes the rest as you.

What you do

39. Read the plan and say what to change; if nothing, say "39 start" and building begins with the experiment site and your songs.

40. Record the twenty-minute voice session on your Mac when you have a quiet room; it is on tkeep and the app's take screen guides it.

38 is still open: ten minutes of clip picks a day, the only way the measure learns your taste moment by moment.

What changed

The whole experiment goes public at cowriter.chrisshaw.me with a simple five-tab interface; your songs, stars, picks and voice sit behind your Google sign-in.

Hit Songs Deconstructed's sheet, about 200 things per song, is now what every song gets; the pilot already reads a third of them, and section 6 lists the rest.

A model of your taste can tell loved from barely kept about 80 % of the time, but 4-stars from 5-stars only about 65 %, so it will say "I don't know" on roughly half the songs rather than guess.

How the co-writer learns and writes What Chris rated, what people replay and his own picks feed one measure of a good song; that measure and a fingerprint of his voice grade his 500 Kismet songs, keep the good parts, fill in the rest or write new songs as him, and his picks, ratings and talk-over notes feed back into the measure. What you rated 2,327 songs 1–5 stars, play counts What people replay clean uploads only comments with a time Your picks two clips, tap one ten minutes a day The measure song grade, moment grade says when it is unsure Your voice and sound range, weight, grit closest singers Your 500 Kismet songs graded, sections ranked good parts kept Fill in · Write like me words, melody, chords into Logic or a free draft You pick, rate, talk over it learns from every answer

1. What it must do

  1. Know what a good song is, and say plainly when it does not know.
  2. Know what is uniquely you: your voice, and the habits in the songs you already wrote.
  3. Grade every one of your songs and every section of it, keep the good parts, and propose the rest.
  4. Write new songs as you, in your range, about what you care about, and get better from your answers.

2. The measure: what a good song is

Kept is the floor

  1. Every song in your library is one you downloaded and kept, so a one-star song still has something you liked; the real bottom is music you never kept.
  2. Four tiers: loved (your 4 and 5 stars), kept (your 1 and 2 stars), kept but never rated (in your library with no stars; the app does not hold them yet, so phase 1 exports them), and not kept (songs from the same years and charts that never entered your library).
  3. The model learns two things separately: kept against not kept, which is your broad taste, and loved against kept, which is your fine taste.
  4. Your 3-stars are treated as "not really rated": they are played far less than 4-stars (a median of 5 plays against 25), so they are left out of training, except the ones you played 20 times or more, which count as a soft 4.

The signals, most trusted first

SignalWhat it measuresWhat distorts it
Your starsYour taste for a whole song, on the 1,827 songs by others you rated.A 3 may mean "never rated"; you rated in different years.
ReplaysWhere everyone rewinds, second by second, on lyric and still-image uploads.The first chorus pulls; music videos are left out.
Timed commentsMoments people name, mostly moments of change.Only about half the songs have enough of them.
Your picksYour taste moment by moment, chorus against chorus.Tiredness, and the edges of the clips.
Play countsMostly repeat your stars.How many years a song has been in your library.
Charts and viewsEveryone's taste for a whole song, the "everyone" grade.Fame, so it is adjusted for the artist's earlier fame.

The checks before anything is trusted

  1. Held-out songs and held-out artists, so the measure cannot simply recognise the band; whole artists are frozen before training and only ever used to test.
  2. It must beat a model that only knows the year, genre and tempo, inside one decade.
  3. The easy test first, loved against kept (562 against 542 songs); then the hard test, 4-stars against 5-stars (377 against 185). If the hard test fails, grades come in three bands, not ten.
  4. Your grade and the "everyone" grade are reported side by side; where they disagree is your taste.
  5. Chance is measured from shuffled stars and shifted replay curves; a difference between two models under 0.07 on the hard test is called a tie, because 185 songs cannot tell finer.
  6. Song grades and moment grades stay separate, and neither is averaged into the other.

3. The taste model: which one, and how sure it is

  1. The song is turned into numbers by a free music model that listened to millions of songs; the first choice is the newest one (MERT, released three days ago), the second is CLAP, which already runs in your app and is forty times faster. Both are used together.
  2. A small model is trained on top of those numbers with your stars: first a simple one, then a slightly bigger one, then the hand-made readings from the pilot added. Retraining the big model itself comes last and only if the simple ones stop improving.
  3. Confidence is made honest three ways: its 80 % must be right 80 % of the time on frozen artists; five copies trained on reshuffled songs must agree; and a song whose sound is far from every training song is marked unknown.
  4. The rule for answering: the model grades a song only when it is at least 80 % sure of the band, the 90 %-sure range is a single band, and the song sounds like songs it trained on. Otherwise it says "I don't know". Expect it to answer 30 to 50 % of songs.
  5. What to expect, from studies that predicted a person's taste from audio: loved against kept about 0.75 to 0.85 (where 0.5 is a coin flip and 1.0 is perfect); 4 against 5 about 0.60 to 0.68. A score of 0.9 is not realistic from audio alone; your picks and notes are what would raise it.
  6. Machine time on the M5 Air, measured today: turning all 1,827 songs into numbers takes about two and a half hours with the newest model and four minutes with CLAP; training and every check take under ten minutes.

4. Your sound and voice

  1. Each hit gets three fingerprints: the voice (a "who is this" set of numbers from a free speaker model), the lead instrument sound (from the listening model already running on your Mac), and the hook melody as steps from the key.
  2. The recognisable part of a song is the four seconds least like any other song that still comes back inside its own song.
  3. A voice is described in plain words: range as notes, weight (deep or light), grit, breath, vibrato, brightness, clear tone, and how the verse voice differs from the chorus voice.
  4. Each of those words is checked on Bono, Michael Jackson, Rob Zombie and Bernard Sumner; a word that ranks them the wrong way round is dropped.
  5. Your session takes twenty minutes on your Mac and nothing is uploaded: slides from low to high, a climbing pattern until it strains, one chorus sung three ways, a minute of talking, one Kismet verse. Your 174 existing takes join it.
  6. Out of it come your natural range and comfortable middle, your tone words, your closest five-star singers, and which of your songs sit in your range and which stretch it.
  7. The co-writer keeps verses in your middle, puts the chorus peak two to four semitones under your top full note, and picks the key so that peak lands there.
  8. The Edge's guitar trick, echoes that land between the picked notes, becomes one reading, so "repeats but does not repeat" can be measured.

5. Attention, familiarity and the eras

  1. The Hey Ya story is real: a Philadelphia station put it between "sticky" songs, Maroon 5 among them, and listeners switching away fell from 27 % to 6 %.
  2. Test one: does a hook that comes back changed beat a plain repeat? Last night found the first hearing wins. It passes only if the lift from change beats chance on held-out songs; otherwise "return the hook" stays the rule and "change it" is dropped.
  3. Test two: do the themes of the year-end hits move from 1960 to 2025 the way you described, and do stacked voices rise after 2020? Year-end charts and lyrics are free; you tag 40 songs by hand so the theme tagger is checked. It passes only if at least four of your six era peaks land where you said.
  4. Test three, the sandwich, needs radio play order, which is not free. The closest test scrapes six iHeart stations' public "recently played" lists for eight weeks as a background job, and only if a one-day trial yields 200 plays. It is a radio fact, not a writing rule.
  5. "Popular with women" cannot be measured before a release; Spotify for Artists shows it for your own releases afterwards, and YouGov names the famous acts women like, which is enough to check whether those acts share traits such as a lower voice, slower tempo and love themes.
  6. The era map is one page, one row per five years since 1960: themes, how deep the lead voice is, stacked voices, tempo, length, solo or band, with your songs plotted on it.
  7. Only tests that pass become rules with numbers. Your 2030s bet (reconnection, independence, relationships questioned, safety, hurt characters, the singer as a person, stacked voices) ships marked as a bet.

6. The song sheet: what Hit Songs Deconstructed tracks

Their free sample report and glossary list about 200 things per song. Every song in the experiment gets the same sheet; the table says how each group is measured and what the pilot already reads.

GroupWhat they trackHow we get it
StructureForm (intro, verse, pre-chorus, chorus, post-chorus, bridge, outro), section lengths and shares, intro length and type, time to the first chorus, chorus count, ending type.The section cutter; missing today.
GenrePrimary genre, sub-genres and influences per section.A free genre tagger; missing today.
InstrumentsNamed instruments, what carries over between sections and what changes, what carries the intro hook.Half read (parts entering and leaving); named instruments need a tagger.
EnergyEnergy within and across sections, first-chorus energy, section lifts (a pull-back, a riser, a drum fill).Read by the pilot.
HarmonyKey, major or minor, chord progressions with their roman numerals per section.A free chord finder; missing today.
MelodyLine letters (A B A B), syllables per line, range, direction, phrase length, motifs, leaps and high notes.Half read (range, held notes); line letters need the section cutter plus the transcript.
Vocal productionLead gender, solo or duet, sung or rapped, doubles, harmonies, ad-libs, vocal pads, effects.Half read (doubles, stacks, noises); gender and delivery need a tagger.
HooksVocal hooks and their dressing, instrumental hooks, tiny hooks, placement, foreshadowing.Half read (parts that return, first hearing); types need the transcript.
Lyrics and rhymesTheme, point of view, title placement and count, imagery and detail, end-of-line rhyme scheme, alliteration, repetition.Half read (rhyme, alliteration, repeats); theme and title need the transcript and a local language model.
Tempo and lengthBeats per minute, song length.Free beat tracker; easy.
BenchmarkHow many of the song's traits match the Top 10's most common value; staying power.Computed once the sheets exist for the year-end hits.

Their stated reason a hit works: "an effective balance of repetition and contrast, or otherwise stated, memorability and engagement", which matches what the pilot found about the first hearing of a returning part.

7. Tools and machine time for the sweep

Every tool was timed today on one real song on the M5 Air, so these are measured hours. All of them are free; the ones with non-commercial weights are used for private grading only.

ToolWhat it givesPer song
Split and fetchVoice, drums, bass, rest; the audio from YouTube.45 s
Section cutterIntro, verse, chorus, bridge, outro with times; beats and tempo.29 s
EssentiaGenre, moods, danceability, energy, arousal and valence, voice or instrumental, bright or dark, key, tempo, loudness.10 s
Beat and chord finderBeats, downbeats, key, major or minor chords.58 s
Note finderEvery note of the voice and the mix as MIDI.10 s
Sound tagger527 sound tags, such as which instruments are heard.3 s
WordsEach sung word with its time.23 s
Sound descriptionsThe numbers the taste model learns from (two models).4 s
The pilot's levelsVoice, instruments, timing, delivery, hooks, lifts, between parts.14 s
Replays and commentsWhere people rewind; comments that name a time.20 s
  1. For 1,827 songs: the graphics-chip jobs (split, words, sound descriptions) take about 37 hours one after another; the processor jobs take about 14 hours with six running at once alongside them; the fetches from YouTube take about 10 hours from any Mac. The whole sweep fits in three days on the M5 Air alone.
  2. The M1 takes the YouTube fetches and half of the processor jobs once it is quiet; today it is overloaded by a backup and cannot be used.
  3. If it slips, drop in this order: the chord finder, then the note finder on the full mix, then the per-second sound tags. Never start the slow pitch trackers or the heavy models.
  4. Stacked voices have no maintained free counter, so the proxy is how many notes sound at once on the voice part plus the stereo width the pilot already reads.
  5. Spotify's audio features are closed to new apps since November 2024; each gets a free stand-in from the tools above, except liveness.

Models that could write or render drafts, tested in this order

  1. ACE-Step 1.5 (free, MIT): writes a full song from lyrics and tags, can repaint a span of seconds, extend, or make a cover from a reference; runs on Apple silicon, about two minutes per minute of song; the best open model on the SongBench test, 6.0 against Suno's 6.9.
  2. YuE2, already installed on your Mac: sings a full song from lyrics and a melody; a three-minute song takes about fifteen minutes on the M5 Air, and a draft mode runs faster than real time.
  3. For MIDI you sing yourself: Composer's Assistant 2 and the Anticipatory Music Transformer, both free, fill in a verse under a given melody.
  4. First test for "fill in a verse under my chorus": ACE-Step repaint with your chorus as the reference. First test for "a full draft from lyrics and a hummed melody": YuE2's melody mode, then ACE-Step with your hummed take as reference.

What earlier projects found

  1. Predicting hits from audio alone tops out around 65 to 75 % right, and the artist's history carries most of the rest; lyrics add about ten percent.
  2. A quarter of streams are skipped in the first five seconds, and skips surge about three and a half seconds after a section boundary, so sections and intros matter.
  3. One published taste model scored 84 % on songs it had seen and 38 to 55 % on new ones, which is why every number here is reported on unseen songs.
  4. Copied into this plan: compare hits with near-misses of the same genre and era, score sections rather than whole songs, and add SongEval, a free musicality grader trained on 2,399 songs rated by sixteen professionals.

8. Your songs and the co-writer

  1. Inventory: 500 Kismet songs in your library, 423 rated 4 or 5. iCloud holds 496 project folders from 2013 on: 35 with a whole-song mix, 284 with only recorded parts, 177 (mostly 2013 sketches) with no plain audio. Songs before 2013 exist only in Apple Music's cloud.
  2. Sections: every song is cut into intro, verse, chorus and bridge by a free structure model, checked against 20 songs you mark by hand, so sections can be judged chorus against chorus.
  3. Grading: the pilot grader, an outside ear that never heard your songs, grades each second; a second grader trained on your stars for others' music says how much a song sounds like what you rate five stars.
  4. Honesty about the outside ear: with the first-chorus pull taken out, its remaining signal on clean uploads is thin (0.07 against 0.06 by chance on 71 songs). So sections get bands, not decimals, until a check on about 900 songs shows it can rank choruses against choruses better than chance.
  5. Your stars on your own songs are never training data; they are memories, and the outside ear must rank your 85 five-star songs above your 338 four-star songs, or the stars are labelled history.
  6. Each song also gets the song sheet from section 6, shown next to the same sheet for the hits, so your songs can be read against hits line by line.
  7. A good part is a section in the song's top band that is unmistakably you: your voice, a melody in your range and typical of your habits, and the most-repeated line heard for the first time. It needs your voice profile first.
  8. Fill in: for a strong chorus and a weak verse, it proposes chords that lead into the chorus, a melody a third under the chorus peak, lyrics in English and Polish, and the rules that held: return the hook, keep every part playing into the chorus, no sudden stop, shorter words.
  9. One Submit, then either MIDI and notes land in Logic or a free audio draft is made with your chorus kept; each draft is a child of the song, so nothing is lost.
  10. Write like me: a third mode on the Create page, pre-filled with your range, tone words, habits and themes; it writes lyrics, melody, chords and arrangement notes, grades each draft before you hear it, and stops at eight drafts or when it beats your median song.
  11. Your stars, your talk-over notes and your picks ("which chorus is more you?") train it, starting with pairs from your own catalogue.

9. The experiment site: cowriter.chrisshaw.me

  1. The whole experiment lives on one site, built the same way as seesound: static pages on Cloudflare, your Google sign-in, and your private data served only to you.
  2. Five tabs: Plan (this page, kept current), Progress (what has been measured, which tools ran, hours used, live), Songs (each song's sheet, grade and confidence), Picks (two clips, tap one; signed in), Findings (what held up, what did not).
  3. Public: the plan, progress, findings, and the sheets for songs by others. Private, behind your sign-in: your 500 songs, your stars, your picks, your voice profile.
  4. The site goes up first, with the plan and an empty progress board, and every phase adds to it, so you can follow the experiment from your phone.

10. Screens, in the order you use them

  1. cowriter.chrisshaw.me: the five tabs above.
  2. Songs list in the app: a grade band beside the stars and a Kismet filter.
  3. Song page "Good parts": sections ranked with one line each, and one Fill-in Submit with a live pipeline.
  4. Project page, Study tab: a grade line over each of the four parts, with the sections marked.
  5. Create: the "Write like me" mode.
  6. Pick: the same screen, with pairs from your own catalogue first.
  7. Voice: the guided session.

Mockups are drawn for all but Pick, then built without waiting for an answer.

11. Build order

PhaseWhat you seeWhere it runs, how long
0cowriter.chrisshaw.me with the plan and a progress board.Cloudflare, like seesound: one day.
1Your 500 songs, and your unrated songs, on the Songs list with their stars.Download on the M1, import on the box: one night.
2Your voice profile and closest singers.Your session on your Mac; the singer test on the M5 Air: a day.
3Every song cut into sections, its sheet filled, and graded second by second.M5 Air, with the M1 alongside: about three days for all 2,327 (section 7 has the hours).
4The taste model with honest confidence on the 1,827 songs by others.M5 Air: half a day.
5Good parts ranked on each of your songs.Needs 2, 3 and 4; then an afternoon.
6Fill in the weak parts, into Logic or a draft.Drafts on the M5 Air where the song model fits, else about an hour of this Mac per song.
7Write like me, and picks that train it.The same, about an hour per song for eight drafts.
8The era map and two attention tests.M1 for the lyric themes, M5 Air for stacked voices: two days, alongside the rest.

Everything runs free on your own machines; nothing in this plan spends money.

12. The biggest risks, each with its cheap check

  1. The outside ear ranks sections no better than chance once the first chorus is taken out: the 900-song check in phase 3 decides bands or decimals.
  2. The taste model learns the era, not the music: compare it with a model that knows only the year and genre.
  3. It recognises artists, not songs: hold out whole artists and watch the score; a drop of a third means it recognised the band.
  4. Your one-star songs are too obscure for replay data: check 50 random one-stars before fetching all.
  5. YouTube blocks a 2,000-song fetch, or serves the wrong version: a 20-song batch first, stop on the first bot check, keep only files within five seconds of the library's length.
  6. The sweep does not fit three days: section 7 names what to drop first.
  7. Your stars on your own songs are memories: the outside grader must rank your 85 five-star Kismet songs above your 338 four-stars, or the stars are labelled history.
  8. Speech voice-prints may not tell singers apart; a past test scored every singer alike against your speech: 20 known singers, two songs each, 90 % must match, or a different model.
  9. The section cutter is wrong on your rougher recordings: 20 songs you mark by hand, and it must agree on 80 % of the boundaries.
  10. The transcriber loses 43 % of your sung words: check five songs; if it stays that bad, you paste lyrics on the Lyrics page.
  11. Rules from other people's hits may not be you: every proposal names its rule, and a rule you reject three times leaves your profile.
Details

Six planners and researchers wrote the sections (the measure, sound and voice, attention and eras, your songs and the co-writer, Hit Songs Deconstructed's criteria, the taste model) and one attacked the merged plan. Their notes are in the music repo under research/cowriter-plan.

Pilot numbers behind the measure: whole-mix readings match replays at 0.154 on unseen songs; adding voice, instruments, hooks and section lifts gives 0.193, chance 0.04; music videos 0.16 against 0.22 on lyric and still-image uploads; the first chorus is the top moment in 39 % of songs against 24 % by chance; with the chorus taken out, 0.07 against 0.06 on the 71 clean uploads.

Your library by stars, songs by others: 168 one-star, 374 two, 723 three, 377 four, 185 five. Kismet: 500 songs, 338 at four, 85 at five, years 2005 to 2024. Only 217 of the songs by others and 25 Kismet songs have audio files on this Mac; the rest are fetched free from YouTube or downloaded from Apple Music. Three-stars: median 5 plays, 13 % never played; four-stars: median 25 plays, 59 % played 20 times or more.

Taste-model sources: the MARBLE benchmark of music models; MERT-v2 (27 September 2026) and MuQ model cards; CLAP (CC0) measured at 0.02 s per ten-second window on the M5 Air; MERT-330M at 0.77 s per 30-second clip; studies of audio-only taste prediction (van den Oord 2013, Lee 2018, Chen 2021) scoring 0.71 to 0.79; a 2026 study showing 0.60 correlation with held-out performers but 0.20 with held-out pieces, which is why artists are held out here; conformal prediction and kNN distance for the "I don't know" rule.

Hit Songs Deconstructed sources: their Cruel Summer sample report (readable without login), glossary, the free 2024 #1 Hit Focus highlights PDF, the Immersion database page and their staying-power articles; ASCAP's 2018 write-up says "over 200 criteria".

Where the work runs: the app's box for import and pages; the M5 Air for splitting, sections, readings, fingerprints and the measure; the M1 for lyric themes and Apple Music downloads; this Mac only for your voice session, and for song drafts if the song model does not fit on the M5 Air (one draft of a three-minute song took about eight minutes here).

U2: "In the Name of the Father" (1993) by Bono, Gavin Friday and Maurice Seezer; Bono sings verses one and four, Friday two and three (Genius, u2songs). The Edge's sound: a dotted-eighth delay (45,000 divided by the tempo, in milliseconds) through two Korg SDD-3000s into two Vox AC30s (amnesta.net/edge_delay; MusicRadar).

Hey Ya: The Power of Habit (Duhigg); WIOQ Philadelphia sandwiched it on 19 September 2003 between 3 Doors Down and Blu Cantrell and on 16 October between Maroon 5 "Harder to Breathe" and Christina Aguilera; switching fell from 26.6 % to 5.7 %.

Free data named in the plan: Billboard year-end lists on Wikipedia and the utdata weekly archive; lyrics from a Kaggle 1959 to 2023 dump and lrclib; MusicBrainz credits; the TidyTuesday Spotify table; YouGov's most popular music artists among women; iHeart stations' public recently-played pages. Paid and therefore out: Mediabase, Luminate, Chartmetric; Spotify editorial playlists left the free API in November 2024.

Standing rules kept: your voice never leaves the Mac and never trains any hosted service; every paid tool needs your yes to an exact amount; no message goes out in your name for a listener panel unless you ask for one.