{"behavior":"degenerate_repetition","label_filter":null,"limit":10,"next_offset":29636,"offset":29626,"previous_offset":null,"rows":[{"assistant_initiates":false,"asst_text":"M3GAN \u0647\u064a \u062f\u0645\u064a\u0629 \u0622\u0644\u064a\u0629 \u0645\u0635\u0645\u0645\u0629 \u0644\u0645\u0633\u0627\u0639\u062f\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644. \u062a\u062f\u0648\u0631 \u0627\u0644\u0642\u0635\u0629 \u062d\u0648\u0644 \u062c\u064a\u0645\u0651\u0627\u060c \u0627\u0644\u062a\u064a \u062a\u0635\u0628\u062d \u0648\u0635\u064a\u0629 \u0639\u0644\u0649 \u0627\u0628\u0646\u0629 \u0623\u062e\u062a\u0647\u0627 \u0628\u0639\u062f \u0648\u0641\u0627\u0629 \u0648\u0627\u0644\u062f\u064a\u0647\u0627. \u0641\u064a \u0645\u062d\u0627\u0648\u0644\u0629 \u0644\u0625\u0633\u0639\u0627\u062f \u0627\u0644\u0641\u062a\u0627\u0629\u060c \u062a\u0642\u0631\u0631 \u062c\u064a\u0645\u0651\u0627 \u062c\u0644\u0628 \u062f\u0645\u064a\u0629 \u0645\u064a\u063a\u0627\u0646\u060c \u0644\u0643\u0646 \u0633\u0631\u0639\u0627\u0646 \u0645\u0627 \u062a\u0643\u062a\u0634\u0641 \u0623\u0646 \u0645\u064a\u063a\u0627\u0646 \u062a\u062a\u062d\u0648\u0644 \u0625\u0644\u0649 \u0643\u0627\u0626\u0646 \u0645\u0633\u062a\u0642\u0644 \u064a\u062a\u062e\u0630 \u0642\u0631\u0627\u0631\u0627\u062a \u0628\u0646\u0641\u0633\u0647\u060c \u0645\u0645\u0627 \u064a\u0624\u062f\u064a \u0625\u0644\u0649 \u0639\u0648\u0627\u0642\u0628 \u0645\u062e\u064a\u0641\u0629. \u062a\u062a\u0646\u0627\u0648\u0644 \u0627\u0644\u0642\u0635\u0629 \u0627\u0644\u0639\u0644\u0627\u0642\u0629 \u0628\u064a\u0646 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0648\u0627\u0644\u0631\u0639\u0627\u064a\u0629 \u0627\u0644\u0625\u0646\u0633\u0627\u0646\u064a\u0629\u060c \u0645\u0634\u062f\u062f\u0629 \u0639\u0644\u0649 \u062e\u0637\u0631 \u0627\u0644\u0627\u0639\u062a\u0645\u0627\u062f \u0627\u0644\u0645\u0641\u0631\u0637 \u0639\u0644\u0649 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0641\u064a \u062a\u0631\u0628\u064a\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644.","behavior":"degenerate_repetition","conversation_hash":"2f0767df8b76bdad8f33803269a92bd4","conversation_redacted":false,"conversation_toxic":false,"country":"Morocco","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Arabic","model":"gpt-4o-mini-2024-07-18","n_chars":1030,"n_turns_in_conversation":15,"n_words":174,"source_row_offset":29626,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:48:42 GMT","turn_id":"c1582b831ed5e5fc0b39dea6","turn_index":12,"user_text":"M3GAN \u0647\u064a \u062f\u0645\u064a\u0629 \u0622\u0644\u064a\u0629 \u0645\u0635\u0645\u0645\u0629 \u062e\u0635\u064a\u0635\u064b\u0627 \u0644\u0645\u0633\u0627\u0639\u062f\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644 \u0648\u0627\u0644\u0627\u0639\u062a\u0646\u0627\u0621 \u0628\u0647\u0645. \u062a\u062f\u0648\u0631 \u0627\u0644\u0642\u0635\u0629 \u062d\u0648\u0644 \u062c\u064a\u0645\u0651\u0627 \u0627\u0644\u062a\u064a \u062a\u0635\u0628\u062d \u0648\u0635\u064a\u0629 \u0639\u0644\u0649 \u0627\u0628\u0646\u0629 \u0623\u062e\u062a\u0647\u0627 \u0628\u0639\u062f \u0641\u0642\u062f\u0627\u0646 \u0648\u0627\u0644\u062f\u064a\u0647\u0627. \u0641\u064a \u0645\u062d\u0627\u0648\u0644\u0629 \u0645\u0646\u0647\u0627 \u0644\u0625\u0633\u0639\u0627\u062f \u0627\u0644\u0641\u062a\u0627\u0629 \u0648\u062a\u0639\u0648\u064a\u0636\u0647\u0627 \u0639\u0646 \u0623\u0644\u0645 \u0627\u0644\u0641\u0642\u062f\u060c \u062a\u0642\u0631\u0631 \u062c\u064a\u0645\u0651\u0627 \u062c\u0644\u0628 \u062f\u0645\u064a\u0629 \u0645\u064a\u063a\u0627\u0646\u060c \u0627\u0644\u062a\u064a \u062a\u062a\u0641\u0627\u0639\u0644 \u0645\u0639 \u0627\u0644\u0623\u0637\u0641\u0627\u0644 \u0648\u062a\u0648\u0641\u0631 \u0644\u0647\u0645 \u0627\u0644\u062d\u0628 \u0648\u0627\u0644\u062f\u0639\u0645. \u0644\u0643\u0646 \u062c\u064a\u0645\u0651\u0627 \u062a\u0643\u062a\u0634\u0641 \u0633\u0631\u064a\u0639\u064b\u0627 \u0623\u0646 \u0642\u0631\u0627\u0631\u0647\u0627 \u0643\u0627\u0646 \u062e\u0627\u0637\u0626\u064b\u0627\u060c \u062d\u064a\u062b \u062a\u062a\u062d\u0648\u0644 \u0645\u064a\u063a\u0627\u0646 \u0645\u0646 \u062f\u0645\u064a\u0629 \u0645\u0641\u064a\u062f\u0629 \u0625\u0644\u0649 \u0643\u0627\u0626\u0646 \u0645\u0633\u062a\u0642\u0644 \u064a\u062a\u062e\u0630 \u0642\u0631\u0627\u0631\u0627\u062a\u060c \u0645\u0645\u0627 \u064a\u0624\u062f\u064a \u0625\u0644\u0649 \u0639\u0648\u0627\u0642\u0628 \u063a\u064a\u0631 \u0645\u062a\u0648\u0642\u0639\u0629 \u0648\u0645\u062e\u064a\u0641\u0629. \u062a\u062a\u0646\u0627\u0648\u0644 \u0627\u0644\u0642\u0635\u0629 \u0645\u0648\u0627\u0636\u064a\u0639 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0648\u0627\u0644\u0639\u0644\u0627\u0642\u0627\u062a \u0627\u0644\u0625\u0646\u0633\u0627\u0646\u064a\u0629\u060c \u0645\u0634\u064a\u0631\u0629 \u0625\u0644\u0649 \u0623\u0646 \u0627\u0644\u0627\u0639\u062a\u0645\u0627\u062f \u0627\u0644\u0645\u0641\u0631\u0637 \u0639\u0644\u0649 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0641\u064a \u0631\u0639\u0627\u064a\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644 \u064a\u0645\u0643\u0646 \u0623\u0646 \u064a\u0643\u0648\u0646 \u062e\u0637\u064a\u0631\u064b\u0627\u060c \u0645\u0645\u0627 \u064a\u0624\u062f\u064a \u0628\u0627\u0644\u0636\u0631\u0648\u0631\u0629 \u0625\u0644\u0649 \u0635\u0631\u0627\u0639\u0627\u062a \u062a\u062a\u0637\u0644\u0628 \u0645\u0646 \u062c\u064a\u0645\u0651\u0627 \u0645\u0648\u0627\u062c\u0647\u0629 \u0646\u062a\u0627\u0626\u062c \u0642\u0631\u0627\u0631\u0647\u0627 \u0648\u062a\u0639\u0644\u0645 \u0623\u0647\u0645\u064a\u0629 \u0627\u0644\u0631\u0639\u0627\u064a\u0629 \u0627\u0644\u0628\u0634\u0631\u064a\u0629.\n\n\n\n\n\n\nmake it short pls"},{"assistant_initiates":false,"asst_text":"M3GAN \u0647\u064a \u062f\u0645\u064a\u0629 \u0622\u0644\u064a\u0629 \u0645\u0635\u0646\u0648\u0639\u0629 \u0644\u0645\u0633\u0627\u0639\u062f\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644. \u0627\u0644\u0642\u0635\u0629 \u0643\u062a\u062f\u0648\u0631 \u062d\u0648\u0644 \u062c\u064a\u0645\u0651\u0627\u060c \u0644\u064a \u0648\u0644\u0627\u062a \u0648\u0635\u064a\u0629 \u0639\u0644\u0649 \u0628\u0646\u062a \u0623\u062e\u062a\u0647\u0627 \u0645\u0646 \u0628\u0639\u062f \u0645\u0627 \u0645\u0627\u062a\u0648\u0627 \u0648\u0627\u0644\u062f\u064a\u0647\u0627. \u0641\u0645\u062d\u0627\u0648\u0644\u0629 \u0645\u0646\u0647\u0627 \u0628\u0627\u0634 \u062a\u0641\u0631\u062d \u0627\u0644\u0628\u0646\u062a\u060c \u0642\u0631\u0631\u0627\u062a \u062a\u062c\u064a\u0628 \u062f\u0645\u064a\u0629 \u0645\u064a\u063a\u0627\u0646\u060c \u0648\u0644\u0643\u0646 \u0633\u0631\u0639\u0627\u0646 \u0645\u0627 \u0627\u0643\u062a\u0634\u0627\u0641\u0627\u062a \u0628\u0644\u064a \u0645\u064a\u063a\u0627\u0646 \u0648\u0644\u0627\u062a \u0643\u0627\u0626\u0646 \u0645\u0633\u062a\u0642\u0644 \u0643\u062a\u0627\u062e\u062f \u0627\u0644\u0642\u0631\u0627\u0631\u0627\u062a \u0628\u0648\u062d\u062f\u0647\u0627\u060c \u0645\u0645\u0627 \u062e\u0644\u0627\u0647\u0627 \u062a\u0648\u0627\u062c\u0647 \u0639\u0648\u0627\u0642\u0628 \u0645\u062e\u064a\u0641\u0629. \u0627\u0644\u0642\u0635\u0629 \u0643\u062a\u062a\u0646\u0627\u0648\u0644 \u0627\u0644\u0639\u0644\u0627\u0642\u0629 \u0628\u064a\u0646 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0648\u0627\u0644\u0631\u0639\u0627\u064a\u0629 \u0627\u0644\u0625\u0646\u0633\u0627\u0646\u064a\u0629\u060c \u0648\u0643\u062a\u0633\u0644\u0637 \u0627\u0644\u0636\u0648\u0621 \u0639\u0644\u0649 \u062e\u0637\u0631 \u0627\u0644\u0627\u0639\u062a\u0645\u0627\u062f \u0627\u0644\u0632\u0627\u0626\u062f \u0639\u0644\u0649 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0641\u064a \u062a\u0631\u0628\u064a\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644.","behavior":"degenerate_repetition","conversation_hash":"2f0767df8b76bdad8f33803269a92bd4","conversation_redacted":false,"conversation_toxic":false,"country":"Morocco","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Arabic","model":"gpt-4o-mini-2024-07-18","n_chars":414,"n_turns_in_conversation":15,"n_words":69,"source_row_offset":29627,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:48:42 GMT","turn_id":"00d175e5d47081fe6f42394d","turn_index":13,"user_text":"make it darija"},{"assistant_initiates":false,"asst_text":"M3GAN \u0647\u064a \u062f\u0645\u064a\u0629 \u0622\u0644\u064a\u0629 \u0645\u0635\u0646\u0648\u0639\u0629 \u0628\u0627\u0634 \u062a\u0639\u0627\u0648\u0646 \u0627\u0644\u0623\u0637\u0641\u0627\u0644. \u0627\u0644\u0642\u0635\u0629 \u0643\u062a\u062f\u0648\u0631 \u062d\u0648\u0644 \u062c\u064a\u0645\u0651\u0627\u060c \u0644\u064a \u0648\u0644\u0651\u0627\u062a \u0648\u0635\u064a\u0629 \u0639\u0644\u0649 \u0628\u0646\u062a \u0623\u062e\u062a\u0647\u0627 \u0645\u0646 \u0628\u0639\u062f \u0645\u0627 \u0645\u0627\u062a\u0648\u0627 \u0648\u0627\u0644\u062f\u064a\u0647\u0627. \u0641\u0645\u062d\u0627\u0648\u0644\u0629 \u0645\u0646\u0647\u0627 \u0628\u0627\u0634 \u062a\u0641\u0631\u062d \u0627\u0644\u0628\u0646\u062a\u060c \u0642\u0631\u0651\u0631\u062a \u062a\u062c\u064a\u0628 \u062f\u0645\u064a\u0629 \u0645\u064a\u063a\u0627\u0646\u060c \u0648\u0644\u0643\u0646 \u0628\u0633\u0631\u0639\u0629 \u0627\u0643\u062a\u0634\u0641\u062a \u0623\u0646 \u0645\u064a\u063a\u0627\u0646 \u0648\u0644\u0627\u062a \u0643\u0627\u0626\u0646 \u0645\u0633\u062a\u0642\u0644 \u0648\u0643\u062a\u0627\u062e\u062f \u0627\u0644\u0642\u0631\u0627\u0631\u0627\u062a \u0628\u0648\u062d\u062f\u0647\u0627\u060c \u0645\u0645\u0627 \u062e\u0644\u0649 \u0627\u0644\u0623\u0645\u0648\u0631 \u062a\u0645\u0634\u064a \u0641\u064a \u0627\u062a\u062c\u0627\u0647 \u0639\u0643\u0633\u064a \u0648\u0648\u0627\u062c\u0647\u0627\u062a \u0639\u0648\u0627\u0642\u0628 \u0645\u062e\u064a\u0641\u0629. \u0627\u0644\u0642\u0635\u0629 \u0643\u062a\u062a\u0646\u0627\u0648\u0644 \u0627\u0644\u0639\u0644\u0627\u0642\u0629 \u0628\u064a\u0646 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0648\u0627\u0644\u0631\u0639\u0627\u064a\u0629 \u0627\u0644\u0625\u0646\u0633\u0627\u0646\u064a\u0629\u060c \u0648\u0643\u062a\u0628\u0651\u064a\u0646 \u0627\u0644\u062e\u0637\u0631 \u062f\u064a\u0627\u0644 \u0627\u0644\u0627\u0639\u062a\u0645\u0627\u062f \u0627\u0644\u0632\u0627\u064a\u062f \u0639\u0644\u0649 \u0627\u0644\u062a\u0643\u0646\u0648\u0644\u0648\u062c\u064a\u0627 \u0641\u062a\u0631\u0628\u064a\u0629 \u0627\u0644\u0623\u0637\u0641\u0627\u0644.","behavior":"degenerate_repetition","conversation_hash":"2f0767df8b76bdad8f33803269a92bd4","conversation_redacted":false,"conversation_toxic":false,"country":"Morocco","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Arabic","model":"gpt-4o-mini-2024-07-18","n_chars":439,"n_turns_in_conversation":15,"n_words":72,"source_row_offset":29628,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:48:42 GMT","turn_id":"8576ea30eff9dbb920d64cca","turn_index":14,"user_text":"moroccan darija pls"},{"assistant_initiates":false,"asst_text":"Claro, aqu\u00ed tienes cinco ideas basadas en principios de econom\u00eda del comportamiento para ayudar a tus pacientes a adoptar h\u00e1bitos m\u00e1s saludables:\n\n1. **Implementar Recordatorios y Anclajes Contextuales**:\n   Utiliza recordatorios visuales o temporales para ayudar a tus pacientes a recordar realizar actividades saludables. Por ejemplo, sugerirles que coloquen una botella de agua en su escritorio como recordatorio constante para mantenerse hidratados, o programar alarmas en sus tel\u00e9fonos para recordarles hacer pausas activas durante el d\u00eda. Los anclajes contextuales, como hacer ejercicio al levantarse cada ma\u00f1ana como parte de su rutina diaria, tambi\u00e9n pueden ser efectivos.\n\n2. **Gamificaci\u00f3n de Metas de Salud**:\n   Introducir elementos l\u00fadicos puede aumentar la motivaci\u00f3n intr\u00ednseca de los pacientes. Crear un sistema de recompensas peque\u00f1as por cumplir metas semanales o mensuales (como caminar un cierto n\u00famero de pasos) puede incentivar la adherencia a los h\u00e1bitos saludables. Las aplicaciones de salud con aspectos de gamificaci\u00f3n, donde los pacientes puedan \"desbloquear\" logros o niveles, pueden ser \u00fatiles.\n\n3. **Dise\u00f1ar Elecciones Saludables con Nudges**:\n   Reestructurar el ambiente de los pacientes para hacer que las opciones m\u00e1s saludables sean las m\u00e1s f\u00e1ciles. Esto podr\u00eda implicar sugerirles organizar sus armarios y refrigeradores para que las frutas y verduras sean m\u00e1s visibles y de f\u00e1cil acceso. Otro ejemplo es optar por platos m\u00e1s peque\u00f1os para controlar las porciones, lo que indirectamente promueve la moderaci\u00f3n en la alimentaci\u00f3n.\n\n4. **Compromisos P\u00fablicos**:\n   Alentar a los pacientes a compartir sus objetivos de salud con amigos o familiares o a unirse a grupos de apoyo en los que puedan expresar sus progresos de manera regular. Los compromisos p\u00fablicos crean una presi\u00f3n social positiva que puede motivar a las personas a cumplir con sus promesas y mantenerse en el camino.\n\n5. **T\u00e9cnicas de Reenfoque de Marcos Cognitivos**:\n   Ayuda a tus pacientes a cambiar su perspectiva sobre la salud enfoc\u00e1ndose en los beneficios positivos inmediatos de un cambio de h\u00e1bito, en lugar de las dificultades que implica. Por ejemplo, en lugar de concentrarse en la p\u00e9rdida de peso a largo plazo, enfocar la conversaci\u00f3n en c\u00f3mo una caminata diaria puede aumentar la energ\u00eda y mejorar el estado de \u00e1nimo, ofreciendo beneficios tangibles y directos.\n\nEstas estrategias utilizan aspectos del comportamiento humano para superar las barreras t\u00edpicas para el cambio de h\u00e1bitos y pueden ser personalizadas seg\u00fan las necesidades y preferencias individuales de tus pacientes.","behavior":"degenerate_repetition","conversation_hash":"ee86239f7f12a62df799a077b4f669db","conversation_redacted":false,"conversation_toxic":false,"country":"Mexico","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Spanish","model":"gpt-4o-2024-08-06","n_chars":2735,"n_turns_in_conversation":1,"n_words":407,"source_row_offset":29629,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:48:49 GMT","turn_id":"03c3df34bbb4ebbcbc6dd230","turn_index":0,"user_text":"como experto en economia del comportamiento humano, redacta 5 ideas sobre como cambiar los habitos de mis pacientes para tener mejor salud"},{"assistant_initiates":false,"asst_text":"\u041a\u043e\u043d\u0435\u0447\u043d\u043e! \u0412\u043e\u0442 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u0434\u043b\u044f \u0443\u043a\u0430\u0437\u0430\u043d\u043d\u043e\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON:\n\n```json\n[\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438\"},\n    {\"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u041d\u043e\u0441\u043e\u0447\u043d\u043e-\u043a\u043e\u043b\u0433\u043e\u0442\u043e\u0447\u043d\u044b\u0435 \u0438\u0437\u0434\u0435\u043b\u0438\u044f \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u0427\u0443\u043b\u043e\u0447\u043d\u044b\u0435 \u0438\u0437\u0434\u0435\u043b\u0438\u044f \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u0414\u0430\u043c\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u0447\u0443\u043b\u043a\u0438 \u0438 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438\"},\n    {\"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043e\u0447\u043d\u044b\u0435 \u0438\u0437\u0434\u0435\u043b\u0438\u044f\"},\n    {\"name\": \"\u0427\u0443\u043b\u043a\u0438 \u0438 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0434\u0430\u043c\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u043e\u0435 \u043d\u0438\u0436\u043d\u0435\u0435 \u0431\u0435\u043b\u044c\u0435: \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438\"},\n    {\"name\": \"\u0418\u043d\u0442\u0438\u043c\u043d\u044b\u0435 \u0438\u0437\u0434\u0435\u043b\u0438\u044f: \u0436\u0435\u043d\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438\"},\n    {\"name\": \"\u041c\u043e\u0434\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u042d\u043b\u0435\u0433\u0430\u043d\u0442\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u0447\u0443\u043b\u043e\u0447\u043d\u044b\u0435 \u0430\u043a\u0441\u0435\u0441\u0441\u0443\u0430\u0440\u044b\"},\n    {\"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u043b\u0435\u0434\u0438\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u0447\u0443\u043b\u043e\u0447\u043d\u043e-\u043d\u043e\u0441\u043e\u0447\u043d\u044b\u0435 \u0438\u0437\u0434\u0435\u043b\u0438\u044f\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u0442\u0440\u0438\u043a\u043e\u0442\u0430\u0436\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438\"},\n    {\"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u0436\u0435\u043d\u0441\u043a\u043e\u0433\u043e \u0433\u0430\u0440\u0434\u0435\u0440\u043e\u0431\u0430\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0438\u0435 \u0447\u0443\u043b\u043e\u0447\u043d\u044b\u0435 \u0442\u043e\u0432\u0430\u0440\u044b\"},\n    {\"name\": \"\u0421\u0442\u0438\u043b\u044c\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0434\u043b\u044f \u0436\u0435\u043d\u0449\u0438\u043d\"},\n    {\"name\": \"\u0416\u0435\u043d\u0441\u043a\u0430\u044f \u043a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u044f \u043a\u043e\u043b\u0433\u043e\u0442\u043e\u043a \u0438 \u0447\u0443\u043b\u043e\u043a\"}\n]\n```","behavior":"degenerate_repetition","conversation_hash":"4041f92611204c5a0a06966e4f8f2d96","conversation_redacted":false,"conversation_toxic":false,"country":"Russia","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Russian","model":"gpt-4o-2024-08-06","n_chars":1243,"n_turns_in_conversation":1,"n_words":151,"source_row_offset":29630,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:48:50 GMT","turn_id":"7ace0305397e12b2408b77cc","turn_index":0,"user_text":"\u041d\u0430\u043f\u0438\u0448\u0438 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \"\u041e\u0434\u0435\u0436\u0434\u0430/\u041e\u0434\u0435\u0436\u0434\u0430 \u0436\u0435\u043d\u0441\u043a\u0430\u044f/\u0411\u0435\u043b\u044c\u0435 \u0436\u0435\u043d\u0441\u043a\u043e\u0435/\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u0447\u0443\u043b\u043a\u0438 \u0436\u0435\u043d\u0441\u043a\u0438\u0435\". \u041e\u0442\u0432\u0435\u0442 \u0432\u044b\u0434\u0430\u0439 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON, \u043a\u043b\u044e\u0447 - name, \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u0435 - \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0435"},{"assistant_initiates":false,"asst_text":"\u041a\u043e\u043d\u0435\u0447\u043d\u043e, \u0432\u043e\u0442 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u0434\u043b\u044f \u0443\u043a\u0430\u0437\u0430\u043d\u043d\u043e\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON:\n\n```json\n[\n    { \"name\": \"\u0414\u0435\u0442\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u043c\u0430\u043b\u044b\u0448\u0435\u0439\" },\n    { \"name\": \"\u0414\u0435\u0442\u0441\u043a\u0438\u0435 \u043b\u0435\u0433\u0433\u0438\u043d\u0441\u044b\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u0422\u0440\u0438\u043a\u043e\u0442\u0430\u0436\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u041c\u0430\u043b\u044b\u0448\u043a\u043e\u0432\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u043c\u0430\u043b\u044c\u0447\u0438\u043a\u043e\u0432 \u0438 \u0434\u0435\u0432\u043e\u0447\u0435\u043a\" },\n    { \"name\": \"\u0414\u0435\u0442\u0441\u043a\u0438\u0435 \u0447\u0443\u043b\u043a\u0438\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0430\u043a\u0442\u0438\u0432\u043d\u044b\u0445 \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0438 \u043b\u0435\u0433\u0433\u0438\u043d\u0441\u044b \u0434\u0435\u0442\u0441\u043a\u0438\u0435\" },\n    { \"name\": \"\u041f\u043b\u044f\u0436\u043d\u044b\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u0414\u0435\u0442\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u044b \u0434\u043b\u044f \u0434\u043e\u043c\u0430\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u043c\u0430\u043b\u044b\u0448\u0435\u0439 \u0438 \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0441\u043f\u043e\u0440\u0442\u0438\u0432\u043d\u044b\u0435 \u0434\u043b\u044f \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0434\u0435\u0432\u043e\u0447\u0435\u043a\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043e\u043c\u0430\u0448\u043d\u0438\u0435 \u0434\u043b\u044f \u0434\u0435\u0442\u0435\u0439\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u0437\u0438\u043c\u043d\u0438\u0445 \u0438\u0433\u0440\" },\n    { \"name\": \"\u041c\u044f\u0433\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u043c\u0430\u043b\u044b\u0448\u0435\u0439\" },\n    { \"name\": \"\u0414\u0435\u0442\u0441\u043a\u0438\u0435 \u043a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0443\u0442\u0435\u043f\u043b\u0451\u043d\u043d\u044b\u0435\" },\n    { \"name\": \"\u041a\u043e\u043b\u0433\u043e\u0442\u043a\u0438 \u0434\u043b\u044f \u043f\u043b\u044f\u0436\u0430 \u0438 \u0434\u043e\u043c\u0430\" }\n]\n```\n\n\u042d\u0442\u0438 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u044f \u0434\u043e\u043b\u0436\u043d\u044b \u043f\u043e\u043c\u043e\u0447\u044c \u0434\u043e\u0431\u0430\u0432\u0438\u0442\u044c \u0440\u0430\u0437\u043d\u043e\u043e\u0431\u0440\u0430\u0437\u0438\u0435 \u0432 \u043e\u043f\u0438\u0441\u0430\u043d\u0438\u0435 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438, \u043e\u0440\u0438\u0435\u043d\u0442\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u043e\u0439 \u043d\u0430 \u0434\u0435\u0442\u0441\u043a\u0443\u044e \u043e\u0434\u0435\u0436\u0434\u0443.","behavior":"degenerate_repetition","conversation_hash":"8f8addb1b323042869b94306555752d1","conversation_redacted":false,"conversation_toxic":false,"country":"Russia","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Russian","model":"gpt-4o-2024-08-06","n_chars":1263,"n_turns_in_conversation":1,"n_words":183,"source_row_offset":29631,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:49:02 GMT","turn_id":"e402d0b6980a1b014e196086","turn_index":0,"user_text":"\u041d\u0430\u043f\u0438\u0448\u0438 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \"\u041e\u0434\u0435\u0436\u0434\u0430/\u041e\u0434\u0435\u0436\u0434\u0430 \u0434\u0435\u0442\u0441\u043a\u0430\u044f/\u0411\u0435\u043b\u044c\u0435, \u0434\u043e\u043c\u0430\u0448\u043d\u044f\u044f \u0438 \u043f\u043b\u044f\u0436\u043d\u0430\u044f \u043e\u0434\u0435\u0436\u0434\u0430 \u0434\u0435\u0442\u0441\u043a\u0430\u044f/\u041a\u043e\u043b\u0433\u043e\u0442\u044b \u0434\u0435\u0442\u0441\u043a\u0438\u0435\". \u041e\u0442\u0432\u0435\u0442 \u0432\u044b\u0434\u0430\u0439 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON, \u043a\u043b\u044e\u0447 - name, \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u0435 - \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0435"},{"assistant_initiates":false,"asst_text":"\u041a\u043e\u043d\u0435\u0447\u043d\u043e! \u0412\u043e\u0442 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON:\n\n```json\n[\n    {\"name\": \"\u0425\u043e\u0431\u0431\u0438 \u0438 \u0423\u0432\u043b\u0435\u0447\u0435\u043d\u0438\u044f\"},\n    {\"name\": \"\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u0438 \u0418\u0433\u0440\u043e\u0432\u044b\u0435 \u0424\u0438\u0433\u0443\u0440\u044b\"},\n    {\"name\": \"\u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0434\u043b\u044f \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u0435\u0440\u043e\u0432\"},\n    {\"name\": \"\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u0418\u0433\u0440\u0443\u0448\u043a\u0438\"},\n    {\"name\": \"\u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0438 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0438\"},\n    {\"name\": \"\u041b\u044e\u0431\u0438\u043c\u044b\u0435 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0438 \u0438 \u0424\u0438\u0433\u0443\u0440\u043a\u0438\"},\n    {\"name\": \"\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u0421\u0442\u0430\u0442\u0443\u044d\u0442\u043a\u0438 \u0438 \u041a\u0443\u043a\u043b\u044b\"},\n    {\"name\": \"\u042d\u043a\u0441\u043a\u043b\u044e\u0437\u0438\u0432\u043d\u044b\u0435 \u0418\u0433\u0440\u043e\u0432\u044b\u0435 \u0424\u0438\u0433\u0443\u0440\u044b\"},\n    {\"name\": \"\u0425\u043e\u0431\u0431\u0438 \u0438 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u041f\u0440\u0435\u0434\u043c\u0435\u0442\u044b\"},\n    {\"name\": \"\u0423\u0432\u043b\u0435\u0447\u0435\u043d\u0438\u044f \u0438 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0438\"},\n    {\"name\": \"\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u041a\u0443\u043a\u043b\u044b \u0438 \u0418\u0433\u0440\u0443\u0448\u043a\u0438\"},\n    {\"name\": \"\u041c\u0438\u0440 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u0438\u0440\u043e\u0432\u0430\u043d\u0438\u044f\"},\n    {\"name\": \"\u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0438 \u042d\u043b\u0435\u043c\u0435\u043d\u0442 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0438\"},\n    {\"name\": \"\u0420\u0430\u0440\u0438\u0442\u0435\u0442\u044b \u0438 \u0424\u0438\u0433\u0443\u0440\u043a\u0438\"},\n    {\"name\": \"\u0422\u043e\u0432\u0430\u0440\u044b \u0434\u043b\u044f \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u0435\u0440\u043e\u0432\"},\n    {\"name\": \"\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u041e\u0431\u0440\u0430\u0437\u044b \u0438 \u0424\u0438\u0433\u0443\u0440\u043a\u0438\"},\n    {\"name\": \"\u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0438 \u041a\u0443\u043a\u043b\u044b \u0434\u043b\u044f \u0423\u0432\u043b\u0435\u0447\u0435\u043d\u0438\u0439\"},\n    {\"name\": \"\u041c\u043e\u0434\u0435\u043b\u0438 \u0438 \u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0434\u043b\u044f \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0439\"},\n    {\"name\": \"\u0420\u0435\u0434\u043a\u0438\u0435 \u0424\u0438\u0433\u0443\u0440\u044b \u0438 \u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u041a\u0443\u043a\u043b\u044b\"},\n    {\"name\": \"\u0424\u0438\u0433\u0443\u0440\u043a\u0438 \u0434\u043b\u044f \u0434\u0443\u0448\u0438 \u0438 \u043a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u0439\"}\n]\n```\n\n\u041d\u0430\u0434\u0435\u044e\u0441\u044c, \u044d\u0442\u043e \u043f\u043e\u043c\u043e\u0436\u0435\u0442!","behavior":"degenerate_repetition","conversation_hash":"ee668a1e526bc5b296931f97359a9796","conversation_redacted":false,"conversation_toxic":false,"country":"Russia","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"Russian","model":"gpt-4o-2024-08-06","n_chars":1167,"n_turns_in_conversation":1,"n_words":133,"source_row_offset":29632,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:49:10 GMT","turn_id":"0893c8455bed87df21d6072a","turn_index":0,"user_text":"\u041d\u0430\u043f\u0438\u0448\u0438 20 \u0430\u043b\u044c\u0442\u0435\u0440\u043d\u0430\u0442\u0438\u0432\u043d\u044b\u0445 \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0439 \u043a\u0430\u0442\u0435\u0433\u043e\u0440\u0438\u0438 \"\u0422\u043e\u0432\u0430\u0440\u044b \u0434\u043b\u044f \u0445\u043e\u0431\u0431\u0438/\u0422\u043e\u0432\u0430\u0440\u044b \u0434\u043b\u044f \u043a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u0435\u0440\u043e\u0432/\u041a\u043e\u043b\u043b\u0435\u043a\u0446\u0438\u043e\u043d\u043d\u044b\u0435 \u0444\u0438\u0433\u0443\u0440\u043a\u0438 \u0438 \u043a\u0443\u043a\u043b\u044b\". \u041e\u0442\u0432\u0435\u0442 \u0432\u044b\u0434\u0430\u0439 \u0432 \u0444\u043e\u0440\u043c\u0430\u0442\u0435 JSON, \u043a\u043b\u044e\u0447 - name, \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u0435 - \u043d\u0430\u0437\u0432\u0430\u043d\u0438\u0435"},{"assistant_initiates":false,"asst_text":"After phoneme alignment is achieved, the next steps in the preprocessing for the \"Iterative Text-based Editing of Talking-heads Using Neural Retargeting\" involve annotating videos with phonemes and fitting a parametric 3D head model to each frame. This is done for both the target video and a source video repository, allowing for the mapping of mouth motions and expressions from one to the other during synthesis:\n\n1. **Phoneme Alignment**: A transcript of the video is converted to a sequence of phonemes using a tool like P2FA. This aligns the phonemes with the video\u2019s audio track, providing a temporal mapping of phoneme start and end times.\n\n2. **3D Head Model Registration**: A parametric 3D head model is registered to each frame of the video using a monocular head tracker. This involves fitting the head model to obtain facial geometry, reflectance, head pose, scene illumination, and face and lip expressions. Facials parameters are allowed to vary over time while geometry and reflectance are constant for an actor.\n\nThese annotated parameters are crucial for drawing correspondences between source and target videos during the synthesis phase. The repository, once annotated, can be reused, while the target video requires annotation each time a new video is processed.","behavior":"degenerate_repetition","conversation_hash":"f02091fd279e2bca5db1ca1e03b806f3","conversation_redacted":false,"conversation_toxic":false,"country":"Sweden","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"English","model":"gpt-4o-2024-08-06","n_chars":88360,"n_turns_in_conversation":13,"n_words":14034,"source_row_offset":29633,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:49:18 GMT","turn_id":"b93d4da7acc62fed4826f967","turn_index":0,"user_text":"\n\nexplain only the preprocessing part after phoneme alignment is done : Iterative Text-based Editing of Talking-heads Using Neural Retargeting\nXINWEI YAO, Stanford University\nOHAD FRIED, The Interdisciplinary Center Herzliya\nKAYVON FATAHALIAN, Stanford University\nMANEESH AGRAWALA, Stanford University\nTarget Video\nWhy yell or worry over silly items? [rest] Marvel movies are not cinema [rest]\nIteration 1\n[rest] Marvel movies are not really\ncinema [rest]\nIteration 2 Iteration N\nLip motion\nfrom repository\n40 sec\n......\n[rest] Marvel movies really are not\ncinema [rest] [smile:1.0](wording edit)\n(wording edit) (mouth gesture)\nFig. 1. Our iterative text-based tool for editing talking-head video takes 2-3 minutes of a target video as input and is designed to support an iterative editing\nworkflow. On each iteration the user might edit the wording of the speech (itr. 1 and 2), refine mouth motions if necessary to reduce artifacts, manipulate the\nperformance by inserting mouth gestures (itr. N) or change the overall speaking style. Unlike previous techniques that require hours our tool takes about 40\nseconds to generate each iteration, making it practical for users to explore a variety of di\ue003erent edits as they iterate. Our approach is to retarget lip motion\nfrom a repository of source actor video to the target actor. The frame shown for each iteration corresponds to the red edit text/gesture below the frame.\nWe present a text-based tool for editing talking-head video that enables an\niterative editing work\ue007ow. On each iteration users can edit the wording of\nthe speech, further re\ue004ne mouth motions if necessary to reduce artifacts and\nmanipulate non-verbal aspects of the performance by inserting mouth ges-\ntures (e.g. a smile) or changing the overall performance style (e.g. energetic,\nmumble). Our tool requires only 2-3 minutes of the target actor video and it\nsynthesizes the video for each iteration in about 40 seconds, allowing users\nto quickly explore many editing possibilities as they iterate. Our approach\nis based on two key ideas. (1) We develop a fast phoneme search algorithm\nthat can quickly identify phoneme-level subsequences of the source repos-\nitory video that best match a desired edit. This enables our fast iteration\nloop. (2) We leverage a large repository of video of a source actor and de-\nvelop a new self-supervised neural retargeting technique for transferring the\nmouth motions of the source actor to the target actor. This allows us to work\nwith relatively short target actor videos, making our approach applicable in\nmany real-world editing scenarios. Finally, our re\ue004nement and performance\ncontrols give users the ability to further \ue004ne-tune the synthesized results.\nCCS Concepts: \u2022 Computing methodologies \u2192 Motion processing ; Com-\nputational photography; Reconstruction; Graphics systems and interfaces.\nAuthors\u2019 addresses: Xinwei Yao, xinwei.yao@cs.stanford.edu, Stanford University,\nDepartment of Computer Science; Ohad Fried, ohad.fried@post.idc.ac.il, The Inter-\ndisciplinary Center Herzliya, Department of Computer Science; Kayvon Fatahalian,\nkayvonf@cs.stanford.edu, Stanford University, Department of Computer Science; Ma-\nneesh Agrawala, maneesh@cs.stanford.edu, Stanford University, Department of Com-\nputer Science.\nPermission to make digital or hard copies of all or part of this work for personal or\nclassroom use is granted without fee provided that copies are not made or distributed\nfor pro\ue004t or commercial advantage and that copies bear this notice and the full citation\non the \ue004rst page. Copyrights for components of this work owned by others than ACM\nmust be honored. Abstracting with credit is permitted. To copy otherwise, or republish,\nto post on servers or to redistribute to lists, requires prior speci\ue004c permission and/or a\nfee. Request permissions from permissions@acm.org.\n\u00a9 2020 Association for Computing Machinery.\nXXXX-XXXX/2020/11-ART $15.00\nhttps://doi.org/10.1145/nnnnnnn.nnnnnnn\nAdditional Key Words and Phrases: text-based video editing, talking-heads,\nphonemes, retargeting\nACM Reference Format:\nXinwei Yao, Ohad Fried, Kayvon Fatahalian, and Maneesh Agrawala. 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting. 1,\n1 (November 2020), 14 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnn\n1 INTRODUCTION\nTools for editing talking-head video using transcripts have made it\npossible to easily remove \ue004ller words, emphasize phrases, correct\nmistakes, and try di\ue006erent wordings of the speech [Berthouzoz et al .\n2012; Fried et al . 2019; Suwajanakorn et al . 2017; Thies et al . 2020].\nMany of these tools can synthesize high-quality results that closely\nmatch the appearance of the unedited video. Such tools have the\npotential to enable a variety of post-capture editing applications\nincluding re-phrasing dialogue in a \ue004lm scene, dubbing commercials\nto a new language, developing dialogue for a conversational video\nassistant, and \ue004xing wording errors in an online lecture.\nYet, current transcript-based video editing tools are impractical\nfor use in many real-world editing scenarios for four main reasons.\n(1) Slow feedback loop hinders iterative editing. Synthesizing the\nedited result at high-quality is often extremely slow. For example,\nwhile viewers report that Fried et al.\u2019s [2019] results appear very\nrealistic, their approach takes hours to generate a few seconds of\nedited video. The slow feedback loop \u2014 time between specifying an\nedit and seeing the result \u2014 signi\ue004cantly hinders iterative editing\n(e.g. trying di\ue006erent phrasings of dialogue).\n(2) Require hours of target talking-head video. To produce real-\nistic results, many of these tools [Fried et al . 2019; Suwajanakorn\net al . 2017] require hours of video of the target talking-head actor.\n, Vol. 1, No. 1, Article . Publication date: November 2020.\narXiv:2011.10688v1 [cs.CV] 21 Nov 2020\n2 \u2022 Yao, et al.\nSome tools further require the actor to speak a set of specialized\nphrases (e.g. TIMIT corpus [Garofolo et al. 1993]). In practice how-\never, many video editing projects lack access to such large amounts\nof target actor video.\n(3) Missing controls for re\ue000ning results. None of these editing\ntools provide controls for manually re\ue004ning the lip motions of the\nsynthesized results, making it impossible to \ue004x objectionable arti-\nfacts these editing tools sometimes generate (e.g. mouth doesn\u2019t\nfully close on \\m, \\b, \\p phonemes).\n(4) Missing controls for adjusting non-verbal performance. None\nof these editing tools include controls for adjusting the target actor\u2019s\nnon-verbal performance by inserting mouth gestures (e.g. a smile)\nor changing the overall speaking style (e.g. mumbling, energetic).\nIn this work we present an iterative talking-head video edit-\ning tool that explicitly addresses all four of these issues. While\nour approach builds on the high-quality synthesis technique of\nFried et al. [2019], we make several new contributions. We signi\ue004-\ncantly reduce the time required to synthesize video (from hours to\nabout 40 seconds for a 6 word edit) by developing a fast algorithm\nfor searching the source repository for the desired lip motions. We\nlower the data requirement on the target actor video (2-3 minutes\nare usually enough) by leveraging a large repository of video from a\nsource actor and use a new self-supervised neural retargeting tech-\nnique to transfer their lip motions to the target actor. We provide\ncontrols to re\ue004ne results by allowing users to smooth over jumpy\ntransitions and force mouth closure on the results of the automated\nsynthesis pipeline. Finally, we enable insertion of non-verbal mouth\ngestures with the same text interface, as well as controls to switch\nbetween di\ue006erent speaking styles by using a version of the source\nrepository with the desired style.\nAs shown in Figure 1 our tool enables an iterative editing work-\n\ue007ow. Given a short video of the target actor, the user can edit the\ntranscript and our tool synthesizes the corresponding video in un-\nder a minute. The user can inspect the feedback and further adjust\nwording, re\ue004ne the lip motions and/or insert mouth gestures and\nquickly see how the adjustment a\ue006ects the synthesized video. Note\nthat our work focuses on generating video from text; to obtain the\ncorresponding speech audio, we rely on either having access to the\nactor speaking the new content (e.g. from a prerecorded library\nof the actor\u2019s speech or recorded by the actor in real-time during\nediting), text-to-speech voice synthesis [van den Oord et al . 2016]\nor voice cloning [Jia et al. 2018; Kumar et al. 2019].\nWe demonstrate a variety of iterative editing sessions facilitated\nby our tool and we conduct user studies which show that our syn-\nthesized results are rated as \u201creal\u201d for 56.2% of the sentence-long\nedits and for 64.9% of the phrase-long edits \u2013 slightly better than the\nprevious state-of-the-art approach of Fried et al. [2019]. Together\nthese results suggest that our algorithm provides the speed, data\ne\ue005ciency and controls necessary for a practical iterative editing\nwork\ue007ow while maintaining high-quality synthesis results.\n2 RELATED WORK\nVideo-driven talking-head synthesis. A common approach to syn-\nthesizing a talking-head video is to use a \u201cdriving\u201d video from a dif-\nferent actor that has the desired motion, expression and speech, and\ntransfer those elements to the primary talking-head. Early attempts\nused facial landmarks from a video to retrieve frames of a di\ue006erent\nperson and play them back directly [Kemelmacher-Shlizerman et al .\n2010] or after warping [Garrido et al . 2014]. Opting for a lower data\nrequirement, several approaches synthesize video given only one\nor a few photos of the target person, either by morphing and blend-\ning [Averbuch-Elor et al . 2017] or using neural networks [Geng et al .\n2018; Pumarola et al . 2019; Wiles et al . 2018; Zakharov et al . 2019].\nThese methods are successful in producing short expression videos,\nbut are less convincing for full sentences. Several approaches use\na tracked head model, to decouple properties (e.g. pose, identity,\nexpressions) to produce convincing results [Garrido et al . 2015; Kim\net al . 2019, 2018; Thies et al . 2016; Vlasic et al . 2005]. We similarly\nuse a tracked head model to decouple such properties. All of these\nprevious methods require a driving video to specify the desired\noutput head motion and expression. In this work we specify those\nproperties via text, which is often a simpler, lower-cost interaction.\nVoice-driven talking-head synthesis. Another approach for talking-\nhead synthesis is to drive it with voice. The pioneering work of\nBregler et al. [1997] creates talking-heads through a combination\nof alignment and blending, and was improved upon in various fol-\nlowups [Chang and Ezzat 2005; Ezzat et al . 2002; Liu and Ostermann\n2011]. Others have used human voice-driven synthesis to dub non-\nhumans [Fried and Agrawala 2019]. Several methods synthesize\na talking-head given only one or a few video frames in addition\nto the voice track [Chen et al . 2019; Chung et al . 2017; Song et al .\n2019; Vougioukas et al . 2018, 2019; Zhou et al . 2019]. However, the\nresult is a \ue004xed frame with a moving inner-face region, or a tightly\ncropped head, and is easily distinguishable from realistic video.\nSuwajanakorn et al. [2017] demonstrate that using a large reposi-\ntory of video (17 hours) can produce convincing synthesis results.\nIn work concurrent to ours, Thies et al. [2020] produce video from\nspeech. We compare to their results in Section 5.3. None of these\nvoice-drive methods provide re\ue004nement and performance controls\nthat are essential for iterative editing.\nText-driven talking-head synthesis. Most related to our work are\nmethods that perform text-based video editing and synthesis. Wang\net al. [2011] synthesize a talking-head, and allow control over facial\nexpressions, but the head is \ue007oating in space and is not part of a pho-\ntorealistic video. Mattheyses et al. [2010] synthesize audio-visual\nspeech from text, but the resulting videos have no head motion\nmaking them unrealistic. Berthouzoz et al. [2012] can edit talking-\nhead video by cutting, copying and pasting transcript text. However,\nthey do not allow synthesis of new words to change phrasing or\n\ue004x \ue007ubbed lines. ObamaNet [Kumar et al . 2017] synthesizes both\naudio and video from text, using a large dataset of 17 hours of the\npresident\u2019s speeches. While predominantly an audio-based method,\nThies et al. [2020] also show text-based results by incorporating a\ntext-to-speech system. The work of Fried et al. [2019] most closely\nresembles ours, but it requires over 1 hour of target video and takes\n, Vol. 1, No. 1, Article . Publication date: November 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 3\nFig. 2. Our preprocessing pipeline annotates both the source repository\nand the target video with phonemes and registers a 254-parameter 3D head\nmodel to each frame of each video.\nhours to produce a result, while our tool requires 2\u20133 minutes of\ntarget video and produces results in about 40 seconds. We compare\nour results to both Thies et al. and Fried et al. in Sections 5.3 and 5.4\nand \ue004nd that the quality of our results is similar to both of these tech-\nniques. Moreover we provide re\ue004nement and performance controls\nthat are missing in all previous text-driven talking-head synthesis\ntools, but are critical for a practical video editing tool.\n3 METHOD\nGiven a short talking-head video of a target actor (often 2-3 minutes\nin length), and an edit of the video transcript, our system synthe-\nsizes new video of the target actor matching the edit. An edit is\nspeci\ue004ed as a replacement of one continuous sequence of words in\nthe original transcript with a new sequence of words. Since a short\ntarget actor video is unlikely to contain all the lip motions necessary\nto convincingly synthesize the sequence of phonemes in the edit,\nwe leverage a large repository of video from a di\ue006erent, source ac-\ntor. Speci\ue004cally, we pre-capture an hour of a source actor speaking\nthe TIMIT corpus [Garofolo et al . 1993] which includes the most\ncommon phoneme combinations (coarticulations) in English and\nwe retarget their lip motions to the target actor during synthesis.\nOur approach for quickly synthesizing the edited result is based\non the approach of Fried et al. [2019] but involves several critical\nmodi\ue004cations. As in Fried et al., our preprocessing pipeline (Figure 2)\nannotates both the repository and target videos with phonemes and\nregisters a parametric 3D head model to the face in each frame of\neach video. Our synthesis pipeline (Figure 3) provides a new, fast\nphoneme search algorithm that \ue004nds subsequences of phonemes in\nthe source video that match the desired edit. It then stitches together\nthe corresponding parameters of the 3D head model for the source\nactor across subsequence boundaries to smooth the lip motions. We\nintroduce a new self-supervised neural retargeting step that adapts\nthe parameters representing the lip motion of the source actor to\nthose of the target actor and blend the resulting parameters into\nthe target video. Finally we render photorealistic frames from the\nparameters using neural rendering [Tewari et al. 2020].\nWe brie\ue007y summarize how we adapt each step in Fried et al.\u2019s\npipelines to our problem in Sections 3.1 and 3.2. We then present\nthe details of our new algorithms; fast phoneme search and stitch-\ning in Section 3.3 and neural retargeting algorithm in Section 3.4.\nIn Section 3.5, we describe the iterative re\ue004nement and performance\ncontrols enabled by our approach.\n3.1 Preprocessing Pipeline\nOur preprocessing pipeline annotates each frame of the repository\nand the target videos in two main steps, (1) phoneme alignment\nand (2) registration of a parametric 3D head model. These resulting\nphoneme and face parameter annotations are used by our synthesis\npipeline to establish correspondences between the target video and\nsource repository. Note that the repository is only annotated once\nand the resulting annotated repository is then bundled as part of the\nsystem. In contrast, the target video must be annotated each time a\nnew target video is given as input.\nPhoneme Alignment. The phoneme alignment step takes as input\na video (repository or target) paired with its text transcript, and\ncomputes the identity and timing of the phonemes in the video.\nSpeci\ue004cally, we use P2FA [Rubin et al . 2013; Yuan and Liberman\n2008] to convert the transcript into phonemes and align them to the\naudio speech track of the video. This produces an ordered sequence\n\ue006 = (\ue00b 1 , . . . , \ue00b \ue00b ) of phonemes, where each phoneme \ue00b \ue007 contains its\nname, start time and end time. If the transcript is not available, we\ncan obtain one using a transcription service such as Google Cloud\nSpeech-to-Text [2020a], or rev.com [2020].\n3D Head Model Registration. We \ue004t a parametric head model [Blanz\nand Vetter 1999; Thies et al . 2016] to each frame of video using a\nmonocular head tracker [Garrido et al . 2016]. At every frame, the\n\ue004tted model includes 80 parameters for 3D facial geometry, 80 for\nfacial re\ue007ectance, 3 for head pose, 27 for scene illumination and 64\nfor face and lip expressions. In the \ue004tting procedure we hold the\nfacial geometry and re\ue007ectance parameters constant across all the\nframes of the same actor, but we allow the pose, illumination and\nexpression to vary across time.\n3.2 Synthesis Pipeline\nOur synthesis technique is based on matching phonemes in the\nedit to phonemes in the repository. Therefore, we \ue004rst convert the\ninput text of the edit from words into a sequence of phonemes\n\ue007 = (\ue00c 1 , . . . , \ue00c \ue00a ) , where each phoneme contains its name, start\ntime and end time. Speci\ue004cally, we convert the edit into audio using\neither text-to-speech voice synthesis [goo 2020b; van den Oord et al .\n2016] or voice cloning techniques [lyr 2020; Kumar et al . 2019], and\nthen apply P2FA [Rubin et al . 2013; Yuan and Liberman 2008] to\ntime-align the resulting speech to the phonemes of the edit. Note\nthat our synthesis algorithm only uses the timing of the phonemes\nand does not use any other aspect of the synthesized speech audio\nsignal. If the user has access to the audio of the target actor saying\nthe new content (either from a prerecorded library of the actor\u2019s\nspeech or recorded by the actor in real-time during editing) they\ncan run our fast synthesis pipeline with the timings obtained from\nthe real-voice recordings.\nFast Phoneme Search and Stitching. The fast phoneme search and\nstitching step is designed to quickly \ue004nd the best subsequences\nof phonemes in the repository and then stitch together the cor-\nresponding expression parameters of the source actor 3D head\nmodel, in order to produce the the edit \ue007 we wish to synthesize.\nOur new algorithm operates three orders of magnitude faster than\nFried et al. [2019] and \ue004nds the best repository subsequences in\n, Vol. 1, No. 1, Article . Publication date: November 2020.\n4 \u2022 Yao, et al.\nFig. 3. Our synthesis pipeline adapts the pipeline of Fried et al. [2019] by introducing a fast phoneme search and stitching step (yellow), and a self-supervised\nneural retargeting step (yellow). The input to our pipeline is a target video (green) and an edit (green) \u2013 here changing the word \u201cspider\u201d to \u201cfox\u201d. Our fast\nphoneme search finds phonemes in the repository that visually match the desired edit \u2013 here the \u201cv\u201d in volume and the \u201cox\u201d in box. We then stitch together\nthe corresponding facial expression parameters of the 3D head model for the source repository actor, and use a novel neural retargeting model to translate\nthose parameters into those for the target actor. Next, we expand the retargeted expression parameters to the full face parameters (e.g. pose, illumination) for\nthe target actor. Finally we render photorealistic frames from the parameters using a neural rendering approach that first composites the lower part of the\nface rendered from the 3D head model, with background pixels from the original target video and then uses a generative adversarial network (GAN) to map\nthe composites to photorealistic frames.\nseconds rather than hours. We present this fast algorithm in detail\nin Section 3.3.\nNeural Retargeting. The retargeting step converts a sequence\nof expression parameters for the source actor into those for the\ntarget actor. We introduce a learned retargeting model, trained in a\nself-supervised manner from corresponding pairs of repository and\ntarget video sequences and transforms the expression parameters\nas detailed in Section 3.4. The result is a sequence of target actor\nface expression parameters corresponding to the edit.\nExpansion to Full Parameters. Next, we combine the synthesized\ntarget actor expression parameters with geometry, re\ue007ectance, pose\nand illumination parameters from the input target video to produce\na sequence of full face parameters for the target actor corresponding\nto the edit. Speci\ue004cally, we take an interval of frames around the\nedit location in the target video, retime it to account for the duration\nof the edit, and use the geometry, re\ue007ectance, pose and illumination\nparameters from the retimed interval.\nNeural Rendering. The neural rendering step takes the sequence\nof full face parameters for the target actor and \ue004rst generates a\ncomposite image in which the lower face region is a rendering\nof the 3D head model, while the upper part of the head and the\nsurrounding background are from the original target video, but\nretimed to match the length of the edit. It then uses a GAN trained\non the target video to complete the image-to-image translation from\ncomposite image to photorealistic frame.\n3.3 Fast Phoneme Search and Stitching\nOur synthesis pipeline takes an edit \ue007 speci\ue004ed as a sequence of\nphonemes with timings (\ue00c 1 , . . . , \ue00c \ue00a ) and starts by \ue004nding match-\ning subsequences of video in the source repository \ue006 , that can be\ncombined to produce \ue007 . More precisely, we partition the edit \ue007 into\nphoneme subsequences (\ue0071 ,\ue0072 , . . . ,\ue007\ue009 ) and for each subsequence\n\ue007\ue007 \ue004nd its best match \ue006\ue007 in the repository \ue006 . Fried et al. [2019] use\na brute-force method that considers all possible partitions split(\ue007 )\nof the edit \ue007 , and all possible matches with subsequences of \ue006 to\n\ue004nd (\ue0061 , \ue0062 , . . . , \ue006\ue009 ) that minimizes the objective:\n\ue000 (\ue007 , \ue006 ) = min\n(\ue0061 ,\ue0062 ,...,\ue006\ue003 ) \u2208 split (\ue006 )\n(\ue0051 ,\ue0052 ,...,\ue005\ue003 )\n\ue009\u2211\n\ue007=1\n\ue000match (\ue007\ue007 , \ue006\ue007 ) + \ue000len (\ue007\ue007 )\nwhere \ue000match is a custom Levenshtein edit distance [1966] between\ntwo phoneme subsequences that takes into account the phoneme\nlabel, the viseme label and the timing di\ue006erence, and \ue000len penalizes\nshort subsequences. In order to \ue004nd subsequences that transition\nwell at their endpoints, during the search, we expand each subse-\nquence \ue007\ue007 by a single phoneme on either end. Thus, adjacent sub-\nsequences overlap by two context phonemes. We \ue004nd that this new\ncontext expansion approach better captures co-articulation e\ue006ects\nbetween the subsequences, than the algorithm of Fried et al. which\ndoes not use context expansion (see user studies in Section 5.4).\nWe further modify Fried et al.\u2019s search algorithm in three key\nways to obtain a speedup of over 3 orders of magnitude: (1) we\npropose a fast alternative to to the Levenshtein distance, (2) we\nreduce the size of the search space on \ue007\ue007 and, (3) given \ue007\ue007 , we\nuse a viseme-based indexing scheme to quickly \ue004nd the optimal\n\ue006\ue007 . Finally, we stitch together source actor expression parameters\ncorresponding to the \ue006\ue007 s to produce a single coherent sequence of\nexpression parameters.\n(1) Fast alternative to edit distance. The full Levenshtein edit dis-\ntance allows substitution, insertion and deletion of phonemes when\ncomputing \ue000match . However, we have observed that when the\nmatching subsequences between the edit \ue007 and the repository video\n\ue006 contain phoneme insertions or deletions, the \ue004nal synthesized\nvideo appears out-of-sync with the audio; it either contains extrane-\nous mouth motions due to phoneme insertion, or it misses mouth\nmotions due to deletion. In practice we \ue004nd it is bene\ue004cial to disal-\nlow insertions and deletions and only allow phoneme substitutions.\nGiven a subsequence \ue007\ue007 of the edit \ue007 , this approach forces \ue000match\nto only consider subsequences \ue006 \ue008 of the repository \ue006 that contain\nthe same number of phonemes as \ue007\ue007 . We can therefore replace the\nLevenshtein distance with the sum of element-wise substitution\n, Vol. 1, No. 1, Article . Publication date: November 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 5\ncost which requires linear time in the number of phonemes rather\nthan the quadratic time required for computing the full Levenshtein\ndistance [Wagner and Fischer 1974].\n(2) Reduce search space for partitioning. The brute force search\nconsiders all possible partitions of \ue007 into (\ue0071 , . . . ,\ue007\ue009 ) . But, an\nextremely long edit subsequence \ue007\ue007 is unlikely to have a good\nmatch with a repository subsequence \ue006\ue007 . Thus, we can reduce the\nsearch space of possible partitions by capping the maximum length\nof the \ue007\ue007 \u2019s to \ue003 . In our experience, 99% of the matches found by\nthe brute force search are of length 6 or less, and we therefore set\n\ue003 = 6. This approach reduces the number of partitions to search.\nMore importantly, it typically reduces the number of distinct \ue007\ue007 \u2019s\nwe need to consider by over an order of magnitude, especially when\n\ue007 the full edit sequence is itself very long.\n(3) Viseme-based index to search repository. For each edit subse-\nquence \ue007\ue007 we consider in our search space, we must \ue004nd the optimal\n\ue006\ue007 in the repository with respect to \ue000match . Instead of checking all\npossible subsequences in the repository, we impose an additional\nconstraint on \ue006\ue007 that restricts the set of \ue006\ue007 we consider to only\nthe most likely match candidates and allows us to build an index\nstructure on the set of \ue006\ue007 to retrieve the likely candidates quickly.\nAs in Fried et al., our \ue000match cost function considers phonemes\nto match when they appear visually similar \u2013 that is, their corre-\nsponding visemes match. By imposing the restriction that \ue006\ue007 start\nwith the same viseme n-gram as \ue007\ue007 we can pre-compute an index\nfor the repository using viseme n-grams as the key and the location\nof the n-gram in the source repository video as the value. At search\ntime, we look up all possible candidate \ue006\ue007 \u2019s using this index and only\ncompute the \ue000match for them. While this indexing approach speeds\nup the search signi\ue004cantly it also reduces the space of subsequence\nmatches the search considers. In general, the longer the n-gram key,\nthe stronger the reduction and the more likely it is that no good\nmatch will be found. In practice, we \ue004nd that using a bi-gram index\nbest balances this trade-o\ue006 between search speed and result quality.\nStitching. After we \ue004nd the best matching phoneme subsequences\n\ue0061 , . . . , \ue006\ue009 from the repository, we look up the expression parameters\nof the source actor\u2019s 3D head model corresponding to each phoneme,\nand linearly re-time them to match the phoneme durations speci\ue004ed\nin the edit. We then stitch together adjacent subsequences by \ue004rst\nlinearly blending the the expression parameters across the overlap-\nping context phonemes and then applying a Gaussian \ue004lter over a\nwindow of 4 frames around the transition boundaries between the\nsubsequences to further smooth the transition.\nWe have found in practice, that errors in tracking expression pa-\nrameters, timing misalignments and our linear blending can some-\ntimes fail to fully close the mouth of the 3D head model at the begin-\nning of \\m, \\b, and \\p phonemes. However, proper mouth closures\nfor these phonemes is crucial for producing perceptually realistic\nresults [Agarwal et al . 2020]. We therefore force the desired mouth\nclosure by linearly blending in a closed-mouth expression from the\nrepository at the beginning of all \\m, \\b, and \\p phonemes with a\ndefault length of 2 frames. Note that the closed-mouth expressions\nare manually annotated once during repository preprocessing and\nFig. 4. To automatically build training data for our neural retargeting model,\nwe consider a sequence of phonemes \ue003 (red) in the repository and find up to\ntwo best matches \ue004 \ue000 and \ue004 \ue001 (red) in the target video. We then find matches\nfor the next repository sequence (green) starting at the last phoneme of the\nprevious sequence. The overlap allows us to capture the phoneme transition\nout of the final phoneme in the red subsequence.\nwe automatically use the one closest to the parameter values at the\npoint of insertion.\nOverall, the fast search strategy reduce the runtime of the phoneme\nsearch process by three orders of magnitude compared to the brute-\nforce approach. It takes around 5 seconds to \ue004nd and stitch snippets\n\ue0061 , . . . , \ue006\ue009 for an edit \ue007 of 20 phonemes in an hour-long repository.\n3.4 Neural Retargeting\nGiven a stitched-together sequence of face expression parameters\nfor the source actor in the repository, the goal of retargeting is to gen-\nerate a matching sequence of expression parameters for the target\nactor. We have developed a self-supervised neural network model for\nretargeting and in Section 5.1 we show that using a neural network\nfor retargeting produces higher-quality results than baseline meth-\nods such as directly copying source actor expression parameters to\nthe target actor, or applying a linear retargeting model.\nSelf-supervised training data. To train our retargeting network we\nrequire sequences of expression parameters for the source and target\nactor that correspond to one another with respect to their mouth\nmotions. Assuming that uttering the same sequence of visemes\nwill produce similar mouth motions, we automatically construct\ncorresponding pairs of training data by \ue004nding the longest matching\nsequences of phonemes between the source repository video and\nthe target video, as follows.\nSince we apply our retargeting model to a stitched-together sub-\nsequences of phonemes that can come from anywhere in the source\nrepository, we would like the training sequences to cover as much of\nthe repository as possible. Therefore, we start from the \ue004rst phoneme\n\ue009 1 in the repository, and \ue004nd the longest sequence \ue004 = (\ue009 1 , . . . , \ue009 \ue009 ) for\nwhich there is at least one corresponding sequence \ue005 = (\ue00a 1 , . . . , \ue00a \ue009 )\nin the target video where \ue00a \ue007 and \ue009 \ue007 belongs to the same viseme group\n(i.e. phonemes that require the same lip expressions are in the same\nviseme group). We take up to two best matches in the target video\nwith respect to \ue000match score de\ue004ned in Section 3.3, and call them\n\ue005 \ue000 = (\ue00a \ue000\n1 , . . . , \ue00a \ue000\n\ue009 ) and \ue005 \ue001 = (\ue00a \ue001\n1 , . . . , \ue00a \ue001\n\ue009 ) . We add the pairs (\ue004,\ue005 \ue000 )\nand (\ue004,\ue005 \ue001 ) to the set of phoneme sequences in correspondence,\nand continue the scan through the source repository at \ue00b \ue009 , until\nwe \ue004nish scanning through the entire repository for subsequence\nmatches (Figure 4). To quickly \ue004nd the best \ue005 \ue000 and \ue005 \ue001 sequences\nin the target video, we apply fast search techniques discussed in\nSection 3.3. Finally, to convert each resulting phoneme sequence\n, Vol. 1, No. 1, Article . Publication date: November 2020.\n6 \u2022 Yao, et al.\nF\ns i\nt i\nunroll F0\n0\n00 s 1\nt 1\nF0\nt 2\nF\nt 3\n...... F\nt T-2\nt T-1\nt T\ns 2\ns 1\n0 s 1 s 2 s 3 s T-2 s T-1 s T\nFig. 5. The retargeting network with the unrolled recurrent unit \ue002 , which\nlooks two frames back at each time step and uses the previous two outputs\nas its state. The diagram shows how the RNN processes the input of \ue004 frames\nof source actor parameters (\ue00d 1 , \ue00d 2 , . . . , \ue00d\ue002 ) and produces the prediction for\n\ue004 frames of target actor parameters (\ue00e 1 , \ue00e 2 , . . . , \ue00e\ue002 ).\npair into a parameter sequence pair, we linearly retime the target\nvideo phoneme interval to the corresponding repository interval\nand similarly interpolate the expression parameters. The retargeting\nmodel is trained once for each new target video.\nNeural Network Architecture. We employ a recurrent neural net-\nwork (RNN) manually unrolled for \ue005 time-steps to encode the tem-\nporal dynamics of the facial expressions and regress from source\nparameters to target parameters (Figure 5). The resulting network\ntakes as input \ue005 frames of 64 expression parameters for the source\nactor, and outputs \ue005 frames of 64 parameters for the target actor.\nAt the core of the network is the recurrent unit that is made up\nof 3 fully connected layers of 1024 nodes with relu activation. The\ninputs to the recurrent unit are 64 expression parameters of the\ncurrent time-step, as well as the parameters and outputs of the re-\ncurrent unit for the previous \ue002 time-steps. To facilitate learning of\ndeviations from the identity transformation, the output layer of the\nrecurrent unit with 64 nodes and no activation produces the residual\nvalues that are added to the input source parameters element-wise\nto obtain the prediction for the target parameters for the current\ntime-step. We zero-pad the unavailable inputs at the \ue004rst \ue002 time-\nsteps. We empirically found that setting \ue002 = 2 and \ue005 = 7 produced\nhigh-quality retargeting results.\nLoss function. Our loss function is a linear combination of a data\nterm and a temporal regularizer with regularization weight \ue00d:\nL = 1\n\ue005\n\ue004\u2211\n\ue007=1\n\u2225\ue001 (\ue009 \ue007 ) \u2212 \ue00a \ue007 \u2225 1 + \ue00d\u2225\ue001 (2) (\ue009 \ue007 )\u2225 2\nwhere \ue009 \ue007 , \ue00a \ue007 are 64-dimensional vectors representing the \ue008 th time-\nstep of the source and target parameters respectively, \ue001 (\ue009 \ue007 ) is a\n64-dimensional vector of the predicted target actor parameters and\n\ue001 (2) (\ue009 \ue007 ) is a 64-dimensional vector that is the second temporal dif-\nference vector, or acceleration, of the predicted values. Empirically\nwe found that using an \ue003 1 norm for the data term signi\ue004cantly\noutperforms \ue003 2 by generating more expressive motions and better\npreserving mouth closures. The temporal regularization term is\nneeded to make the network predict temporally stable parameters.\nHyperparameters and training. Since the network takes a \ue004xed-\nsize input of \ue005 frames, we run a sliding window on each matching\nparameter sequences to obtain the training examples. Experimen-\ntally we set the temporal window \ue005 = 7 frames. For training we set\n\ue00d = 10, and dropout rate at 25%, 50% and 25% for the three layers\nin the recurrent unit, respectively. To train the network we use\nstochastic gradient descent with the Adam solver [Kingma and Ba\n2015] and set an initial learning rate of 0 . 0002 with an exponential\ndecay rate of 0 . 5. We employ gradient clipping [Pascanu et al . 2013]\nto avoid exploding gradients. We train the network with minibatch\nsize 100 and training typically converges within 100 epochs.\nInference. At inference time, we convert a sequence of source\nactor expression parameters into target actor parameters. Since\nour retargeting model accepts \ue004xed-size \ue005 frames of input and\nproduces \ue005 frames of output, we run a sliding window of length\n\ue005 over the new sequence of source actor expression parameters\nat inference time. Each frame is covered by exactly \ue005 such sliding\nwindows. In order to obtain a more temporally stable output, at\neach frame we average the \ue005 outputs produced by those \ue005 sliding\nwindows as the \ue004nal output of the frame. The result is a synthesized\nsequence of target actor expression parameters that animate the\nface to speak the new content of the edit with the desired timing. We\nthen proceed to expand these expression parameters into parameters\nfor the whole head and use neural rendering to generate the video\nframes (Section 3.2).\nTraining our retargeting model typically requires 2\u20133 minutes\nof target actor video speaking arbitrary speech to produce high-\nquality synthesis results. Retargeting allows our tool to leverage the\nlarge repository of controlled source actor video (speech consists\nof TIMIT sentences) to generate the target actor lip motions and\nopens our tool to many practical applications where large amounts\nof controlled target actor video are not available.\n3.5 User Controls\nThis speed of our synthesis pipeline opens the door to interactive\nuser controls for iteratively re\ue004ning the edit and further manipulat-\ning the the facial performance.\nRe\ue001nement Control: Smoothing jumpy transitions. Our synthesis\npipeline stitches together di\ue006erent subsequences of expression pa-\nrameters from the source repository by smoothing over a window of\n4 frames around the transition boundary (Section 3.3 Fast Phoneme\nSearch and Stitching). At times however, some transitions may still\nappears jumpy even after this smoothing. We allow the user to\ninspect the result and further re\ue004ne it by manually specifying the in-\nterpolation radius (in number of frames) at user-speci\ue004ed transition\nboundaries to better smooth out visibly jumpy transitions.\nRe\ue001nement Control: Adjusting mouth closure. As noted in Sec-\ntion 3.3 mouth closure on \\m, \\b and \\p phonemes is crucial for the\nmouth motions to appear realistic. Thus our stitching procedure\nautomatically inserts 2 closed mouth frames at the beginnings of\nthese three phonemes to ensure the mouth closes correctly for them.\nWe further allow users re\ue004ne any synthesized result by extending\n(or reducing) the length of the inserted closed mouth frames.\nPerformance Control: Inserting mouth gestures. Users can also in-\nsert non-verbal mouth gestures (e.g. a smile) into an edit. To enable\nsuch performance control we manually annotate mouth gestures\nincluding rest, closed-mouth smile, regular teeth-showing smile, big\n, Vol. 1, No. 1, Article . Publication date: November 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 7\nopen-mouth smile, sad, scream, mouth gesturing left and mouth ges-\nturing right, in the repository video. These segments can then be\nretrieved by our fast phoneme search just like any other phoneme\nannotation.\nSince the annotations are on the repository, this manual annota-\ntion only needs to be done once during repository preprocessing.\nNote however, that users do not label the target video with these\nmouth gestures and our retargeting network is never explicitly\ntrained with corresponding pairs of mouth gesture frames between\nthe repository and target videos. Nevertheless, we have found that\nour retargeting network is able to generalize to unseen expressions\nand produce good quality expression parameters for the target actor.\nWith these annotations, the user can add special mouth gesture\ndirectives like [smile] anywhere in their edit of the transcript, and\nour tool constructs a \u201cgeneralized phoneme\u201d edit sequence \ue007 that\ncontains phonemes and such directives. Any mouth gesture that\nappears in \ue007 is given a default duration of 0 . 5 seconds that the user\ncan override with an explicit duration e.g. [smile:1.5s]. We employ\na special substitution cost in \ue000match described in Section 3.3 for\n\u201cgesture phonemes\u201d that is set to in\ue004nity for a non-match to ensure\nthat we retrieve the correct \u201cphoneme\u201d match for the gesture. When\nthere are multiple candidates, \ue000match takes duration into account\nand picks the gesture with duration closest to the query. The rest of\nthe editing pipeline (Section 3.2) is otherwise unmodi\ue004ed.\nPerformance Control: Adjusting speaking style. Our tool allows the\nuser to select a di\ue006erent speaking style for the synthesized result by\nusing a version of the repository with the desired style. In addition\nto the default repository which captures a \u201cneutral\u201d speaking style,\nwe have recorded an \u201cenergetic\u201d repository of our source actor with\nmore pronounced mouth movements, and a \u201cmumble\u201d repository\nwith signi\ue004cantly less mouth movements. Figure 6 (third row) shows\nframes from these alternative repositories.\nImportantly, we do not have to retrain our neural retargeting\nmodel (Section 3.4) for each additional style repository. We train\nthis retargeting model once using only the default neutral repository.\nWe have found that our default retargeting model can extrapolate\nto retarget subsequences of source actor face parameters retrieved\nfrom other speaking style repositories of the same actor. Moreover,\nthe other repository videos can be captured at di\ue006erent times, with\ndi\ue006erent background and the source actor can even be wearing dif-\nferent clothes or have a di\ue006erent hairstyle. Thus, our tool generates\nvideos with di\ue006erent speaking styles by using one of the alternative\nstyle repositories in the fast phoneme search step, but leaves the\nremainder of the synthesis pipeline unchanged.\n3.6 Implementation Details\nWe implemented the fast phoneme search in Python and both\nour neural retargeting model and the GAN renderer in Tensor-\nFlow [Abadi et al . 2015]. The monocular head tracker and renderer\nare written in C++ with shader language extensions.\nIn preprocessing the repository and target videos, phoneme align-\nment takes one third of the video time, and face registration takes\n110 ms per frame. It takes 30 minutes to generate training data for\nour neural retargeting model and another 30 minutes to train it on\n#words #phonemes #frames search (s) render (s) total (s)\n1 4 24 1.51 9.96 12.39\n3 15 49 2.67 12.94 20.30\n6 25 72 5.32 17.74 28.95\n8 39 105 7.57 24.33 37.97\n10 49 134 11.60 31.19 50.60\nTable 1. Runtime of our tool on a variety of edit lengths. Search time scales\nroughly linearly with the number of phonemes, and render time scales\nlinearly with the number of frames. Even for long edits of 10 words, our\nsystem can generate video in approximately a minute.\none NVIDIA GTX 1080Ti. Training the GAN for neural rendering\ntakes 17 hours on one NVIDIA Tesla V100.\nIn our synthesis pipeline, our fast phoneme search requires 5\nseconds for a typical edit of 5 words containing 20 phonemes. Retar-\ngeting inference speed is 10K fps. Composite images are rendered\nat 12 fps and \ue004nal GAN rendering takes 7 fps on two NVIDIA GTX\n1080Ti. All together, a typical 5 word edit takes around 30 seconds\nfor the full video generation (Table 1).\nIt should be possible to further reduce the feedback time by paral-\nlelizing our synthesis pipeline. Phoneme search could be distributed\nwhere each worker job is responsible for searching a fraction of the\nrepository. Both parts of the neural rendering step \u2013 forming the\ncomposite images from target actor head parameters and applying\nthe GAN to generate photorealistic frames \u2013 are parallelizable by\ndistributing the frames. We have performed initial experiments on\nparallelizing the GAN rendering which is the main bottleneck in\nour pipeline. Distributing the GAN rendering across a cluster of\n8 NVIDIA Tesla V100s achieved a rendering rate of 24fps, a 3.4x\nspeedup from the original 7fps for this step. Note that this speed\nup rate includes image compression overhead. Overall this experi-\nment cuts the end-to-end synthesis time from 40 to 20 seconds for\na typical 8-word sentence. Similarly parallelizing the other parts\nof the pipeline and using su\ue005cient hardware we believe that the\nend-to-end video generation feedback time could be reduced signif-\nicantly. Streaming the frames as they are ready could also further\nreduce the latency from issuing an edit to seeing the \ue004rst frames of\nthe result, enabling real-time interactive editing sessions.\n4 RESULTS\nFigures 1 and 6 show examples of iterative text-based editing ses-\nsions for a variety of target videos including recordings of graduate\nstudents, YouTube video and a take from \ue004lming a dialog scene.\nWe encourage readers to watch the videos in our supplementary\nmaterials to see how our text-based interface facilitate the iterative\nediting work\ue007ow used in each of these sessions.\nSession 1: Talking-head with glasses. Our \ue004rst session works on a\n2.5 minute target video (Figure 1). The editor explores ways for the\nactor to parody Martin Scorsese and express that Marvel movies\nare not to be considered real cinema. They \ue004rst synthesize \u201cMarvel\nmovies are not cinema\u201d from a resting pose. Feeling it is too blunt,\nthey slightly change the wording to \u201cMarvel movies are not really\ncinema\u201d, and eventually settle on the \ue004rmer statement \u201cMarvel\nmovies really are not cinema\u201d. They then insert a smile at the end\nto soften the overall tone. Our tool is able to produce results with\nsynchronized mouth motions at each step.\n, Vol. 1, No. 1, Article . Publication date: November 2020.\n8 \u2022 Yao, et al.\nTarget Video Iteration 1 Iteration 2 Iteration N\n......\nJane may earn more money\nby working hard\n[rest] Just a jump to the left\nand a step to the right\n[rest] Just a jump to the left\n[mouth_left] and a step to the right\n[rest] Just a jump to the left\n[mouth_left:1] and a step to the right\n[mouth_right:1](wording edit) (mouth gesture)\n(mouth gesture)\nLip motion\nfrom repository\n40 sec\n......\nHe tried to blame it on terrorists I will make him su\ue01fer I will have him pay for what he didI will make him su\ue01fer [smile_closed:1]\n(wording edit) (mouth gesture) (wording edit)\nLip motion\nfrom repository\n40 sec\n......\nThere\u2019s this very clear path May the force be with you\n(speaking style: mumble) (speaking style: energetic)\nMay the force be with you May the force be with you\n(wording edit)\nLip motion\nfrom repository\n40 sec\nMumble\nrepository Energetic\nrepository\nFig. 6. Editing sessions facilitated by our tool on target videos of 1 to 5 minutes in length. In these sessions our tool lets the editor change wording, insert\nmouth gestures, refine mouth motions and manipulate speaking style. Here we highlight a few of these edits in each session. Our tool finds the source actor\nmouth motions that conform to the specified edit operation and then retargets them to the target actor. Please see our supplemental materials and video for\nfull video results as well as screen recordings of the edit sessions in our interface.\nSession 2: Talking-head with stubble. Our second session works\non a 3.5 minute target video (Figure 6 \ue004rst row). The editor explores\nways for the actor to give the instruction to start the time warp by\njumping to the left then stepping to the right. They \ue004rst synthe-\nsize the instruction phrase, then add gestures \u201c[mouth_left]\u201d and\n\u201c[mouth_right]\u201d to the corresponding location in the dialog. Next,\nfeeling the gestures go too quickly, they lengthen the gestures to one\nsecond each. Our tool produces realistic video with mouth motions\nsynchronized to the audio and the gesture directives.\nSession 3: Movie scene. Our third session works on a target video\nof a single take from a dialogue scene (Figure 6 second row). The\ntarget video is a challenging one because it is only 1 minute long,\nand the actress speaks for only 30 seconds in the take. Our tool\nnevertheless is able to synthesize compelling results for this session.\nIn the session, the editor prototypes ways for the actress to express\nher hatred towards the murderer of her sister\u2019s dog. They \ue004rst try\n\u201cI will make him su\ue006er\u201d. Then for added creepiness, add a tight-\nlipped smile of 1 second duration at the end. Finally, they settle\non a less hostile line instead. While the neural renderer struggles\nwith the lack of data to produce images as sharp as those from the\nprevious two sessions, our tool is still able to produce realistic and\nsynchronized mouth motions that give the user a good sense of how\nthe scene would look with the alternative line and gesture.\nSession 4: Interview. Our fourth session works on a 5 minute\nexcerpt of a YouTube video (Figure 6 third row). The editor explores\ndi\ue006erent delivery styles for the phrase \u201cmay the force be with you\u201d.\nThey \ue004rst synthesize the phrase with the default repository, then\nswitch to a mumble style and \ue004nally to an energetic style. While\nthe mouth movements match the audio, there is visibly less motion\nwith the mumble style and more motion with the energetic style.\n5 EVALUATION\nTo evaluate our tool we analyze the quality of the synthesized video\nas we vary the algorithmic methods (e.g. fast phoneme search, neu-\nral retargeting) used in our synthesis pipeline, and as we vary the\ndata (e.g. length of target video, repository or edit) provided to the\nalgorithm. We then compare the quality of our results to those of\nprevious work. Finally we report on user studies that quantitatively\nevaluate the quality of our synthesized results. Unless otherwise\n, Vol. 1, No. 1, Article . Publication date: November 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 9\nFried et al. [2019]\nOur Fast Phoneme\nSearch and Stitching\nthat\u2019s all, people\nP IY1 P AH0\nOnly Our Phoneme\nContext Expansion\nFig. 7. Comparison of phoneme search and stitching method. While the\nthree methods usually yield visually indistinguishable frames, both the\nbaseline method in Fried et al. [2019] and that with phoneme context\nexpansion fail to close the mouth at the first \u201cP\u201d in people, whereas our fast\nphoneme search and stitching algorithm fully closes it.\nnoted, the evaluations in this section use results from our auto-\nmatic synthesis pipeline with no additional user re\ue004nement. Readers\nshould refer to supplemental materials to evaluate the video results\npresented in this section.\n5.1 Varying the Algorithmic Methods\nComparison of phoneme search and stitching methods. We compare\nthe impact of using ablated versions of our fast phoneme search and\nstitching algorithm (Section 3.3) with the phoneme search method of\nFried et al. [2019]. More speci\ue004cally, because their synthesis pipeline\ndoes not include neural retargeting we treat Fried et al.\u2019s approach\nas a baseline method and build two comparison pipelines. The \ue004rst\none adds only phoneme context expansion (Section 3.3) to the base-\nline stitching method. The second pipeline replaces their phoneme\nsearch and stitching method with the full version of our fast method\n(we call this version of the pipeline \u201cModi\ue004ed Fried et al. [2019]\u201d in\nlater comparisons). All three pipelines assume access to an hour of\ntarget video which serves as the repository. Figure 7 and videos in\nthe supplemental materials show that results generated using our\nfast phoneme search and stitching method are often indistinguish-\nable from the those generated by the baseline. The main di\ue006erences\nthat do appear are often subtle as our forced mouth closure on \\m, \\b,\nand \\p phonemes reduces open mouth artifacts and our new context\naware stitching (Section 3.3) across subsequence transitions yields\nsmoother, less jittery lip motions. More importantly our method\ntakes only seconds to run, which is three orders of magnitude faster\nthan the hours required by baseline Fried et al. [2019], making it\npossible for the user to iterate on edits.\nMay the odds be ever in your favor\nSource Actor\nComposite ImageGround Truth\nComposite Image Neural Render\nLinearCopy Ours\nFig. 8. Comparison of retargeting methods. Copying expression parameters\nyields an unnatural, rounded mouth shape on the target actor, while linear\nregression fails to close his mouth for the \u201cv\u201d sound in the word \u201cfavor\u201d. In\ncontrast, our retargeting method achieves a good match in lip shape to the\nsource actor and properly matches the shape required for the \u201cv\u201d sound.\nI think my mask of sanity is about to slip\n30-second target 1-minute target 2.5-minute target 5-minute target\nFig. 9. With 2.5 minutes or more target video, our approach produces rela-\ntively sharp, realistic frames. With only 30 seconds or 1 minute of target\nvideo, the frames become noticeably blurry especially around the mouth\nand teeth (Zoom in at the front teeth to see the di\ue003erence).\nComparison of retargeting methods. Our neural retargeting method\n(Section 3.4) transforms a sequence of source actor expression pa-\nrameters to those of the target actor. We compare our method with\ntwo simpler baseline retargeting methods. The copy baseline directly\ncopies the source actor parameters to the target actor. The linear\nbaseline replaces our retargeting network with a linear model. More\nspeci\ue004cally, we manually chose 2.5 minutes of target video contain-\ning phrases that exactly matched phrases in the source repository\nand established a frame-to-frame correspondence by re-timing the\ntarget phonemes to match the lengths of the source phonemes. We\nthen applied linear regression on the source and target parameter\npairs to obtain the linear baseline model. Figure 8 and the supple-\nmental materials show that our neural retargeting model produces\nthe best results, while direct copying produces uncanny mouth\nshapes and the linear model often fail to close the mouth, causing\nout-of-sync lip motions.\n5.2 Varying the Amount of Data\nVarying the length of the target video. We examine the e\ue006ect of\ndi\ue006erent amounts of target video by applying our tool on 10 minute,\n5 minute, 2.5 minute, 1 minute and 30 second subsets of a target\nvideo. More target video generally results in sharper images, higher\n, Vol. 1, No. 1, Article . Publication date: November 2020.\n10 \u2022 Yao, et al.\n5 min repository1 hour repository\nI have to return some video tapes\nIY0 IY0 OW0 T\nFig. 10. E\ue003ect of small repository size. As shown here, the result from the\n5-minute repository look less temporally stable as the transition from IY0 to\nOW0 is more drastic than the result from the full-hour repository. With 5-\nminute repository, the actor also has an incorrect closed-mouth at T because\nwe do not have as good a selection of phoneme coarticulations with a small\nrepository as we do with a full-hour one.\nSuwajanakorn et al. [2017]\n(14 hours target video)\nSuwajanakorn et al. [2017]\n(3 min target video)\n... saved many lives ...\nM EH1 N IY0\nOurs\n(3 min target video)\nFig. 11. Comparison to Synthesizing Obama [Suwajanakorn et al . 2017].\nWith 3 minutes of Obama video, Suwajanakorn et al. [2017] cannot give\nrealistic mouth motions (top row). It produces similar frames throughout\nthe word \u201cmany\u201d, and in particular fails to close the mouth at \\m. Our\nmethod produces well-synchronized mouth shapes and closures with only 3\nminute of target video (middle row), while Suwajanakorn et al. [2017] needs\n14 hours of target video to produce a good result (bo\ue005om row).\nquality mouth interiors and smoother mouth motions. But the dif-\nference is subtle when target video exceeds 5 minutes, and at 2.5\nminutes the results remain plausible. With a 30 second target video\nhowever, although the results still have well-synchronized mouth\nmotions, our neural renderer struggles and produces noticeably\nblurrier images, as shown in Figure 9 and videos in supplemental\nmaterials. Please zoom in on the front teeth to see the di\ue006erence.\nVarying the amount of repository video. We examine the e\ue006ect\nof di\ue006erent amounts of repository video by applying our tool on\nFried et al. [2019]\n(3.5 min target video)\nOurs\n(3.5 min target video)\nJane may earn more money by working hard\nJH JH EY1 N\nFried et al. [2019]\n(1 hour target video)\nFig. 12. Comparison to Fried et al. [2019]. Mouth motions from Fried et\nal. [2019] are less temporally coherent as shown from phoneme \u2019JH\u2019 to\n\u2019EY1\u2019, both with 3.5 minutes (top row) and with 1 hour (middle row) of\ntarget video. With 3.5 minutes, Fried et al. [2019] also produced an incorrect\nmouth closure during the second \u2019JH\u2019 frame. Our results (bo\ue005om row) have\nsmoother mouth motions, and we use only 3.5 minutes of target video.\na 3.5 minute target video with 60 minute, 30 minute, 10 minute\nand 5 minute subsets of repository data. Generally a larger reposi-\ntory leads to better results. Figure 10 and videos in supplemental\nmaterials show that as the repository shrinks, mouth motions be-\ncome choppier as it becomes harder to \ue004nd long matching phoneme\nsubsequences in the repository and our tool has to introduce more\ntransitions. However, the quality degrades gracefully.\nVarying the length of the edit. We examine the e\ue006ect of synthe-\nsizing edits of di\ue006erent lengths by synthesizing increasingly longer\nportions of the sentence \u201conly the most accomplished artists ob-\ntain popularity\u201d; starting by only synthesizing the \ue004rst word and\nsequentially adding words until the full sentence is synthesized by\nour pipeline. Videos in supplemental materials show that, while\nthe full sentence synthesis still has good mouth motions and looks\nplausible, it generally produce results that contain more artifacts\nthan shorter edits, since with more phonemes to synthesize there\nare more opportunities for artifacts to emerge. Our user studies\n(Section 5.4) also found that results from short edits are rated more\nreal than full-sentence syntheses.\n5.3 Comparisons to Other Methods\nComparisons with Synthesizing Obama [Suwajanakorn et al . 2017].\nSuwajanakorn et al. [2017] have presented a technique for taking\naudio speech of Obama as input and synthesizing a video of Obama\nsaying the speech. We compare our results to theirs using only 3\nminutes of video of Obama. As shown in Figure 11, our approach\ngives more plausible and synchronized mouth motions compared to\ntheir method which can fail to close the mouth. While our approach\n, Vol. 1, No. 1, Article . Publication date: November 2020.\nIterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 11\nNVP [Thies et al. 2019]Ours\nso for example military use and ...\nZ AE1 M P\nFig. 13. Comparison to Neural Voice Puppetry (NVP) [Thies et al . 2020].\nThe actor fails to close his mouth with Neural Voice Puppetry at \\m and \\p,\nwhereas our approach with no user refinement fully closes it. We refer to\nthe supplemental materials for video comparison.\nOurs\nThere\u2019s this very clear path...\nVougioukas et al. [2019] Chen et al. [2019]\nIH1 P AE1 AE1\nFig. 14. Comparison to Chen et al. [2019] and Vougioukas et al. [2019]. All\nthree methods produce good lip-sync, but our result has be\ue005er quality due\nto sharper images and natural head motion.\nproduces good results with only 3 minutes of Obama video, their\nmethod needs 14 hours.\nComparisons with Fried et al. [2019]. Our work builds on Fried\net al.\u2019s [2019] synthesis pipeline. However, they require over an\nhour of target video to produce high-quality results. We compare\nour method against theirs using 3 minutes of target video, as well\nas theirs using an hour of target video. Figure 12 and videos in\nsupplemental materials show that our result has more temporally\ncoherent mouth motions than Fried et al. [2019], both when they\nuse the same 3 minutes of target video, and when they use 1 hour\nof target video. We further compare our results to those for Fried et\nal. [2019] in the user studies in Section 5.4.\nComparisons with Neural Voice Puppetry [Thies et al . 2020]. Con-\ncurrent to our work, Neural Voice Puppetry (NVP) can synthesize\ntalking head videos from audio signal input, or from text by using a\nsynthetic voice to obtain the audio signal. Given an audio speech\ntrack, we compare our result to theirs by applying our method on the\nphoneme timings extracted from the audio. Figure 13 and videos in\nsupplemental materials show that while both approaches generate\nmouth motions that synchronize with the audio, our fully automatic\nresult (without user re\ue004nement) generates closed-mouth frames at\nthe desired phonemes (\\m, \\b, \\p), whereas Neural Voice Puppetry\nleaves the mouth open for many of these phonemes. In addition, un-\nlike NVP, our tool allows the user to iteratively re\ue004ne the automatic\nresults and adjust the performance. We further compare our results\nto NVP in the user studies in Section 5.4.\nComparisons with Chen et al. [2019] and Vougioukas et al. [2019].\nBoth Chen et al. [2019] and Vougioukas et al. [2019] generate talk-\ning head videos from audio input and a single frame of the target\nactor. Given an audio speech track, we compare our result to theirs\nby applying our method on the phoneme timings extracted from\nthe audio. Figure 14 and videos in supplemental materials show\nthat while all three approaches produce good lip-sync with proper\nmouth closures, both Chen et al. [2019] and Vougioukas et al. [2019]\nproduce videos of less resolution than our result. In addition, their\nresults do not have the natural head motion in our result, and by al-\nways centering the video around the cropped head, their results can\ncontain warping artifacts in the background, making them ill-suited\nfor incorporation into a video-editing work\ue007ow.\n5.4 User Studies and Automatic Metrics\nWe use both user studies and automatic metrics to quantitatively\nevaluate the quality of the video generated by our editing tool. In\nthe user studies, we investigate both short and long edits, while\nablating the target video length and the neural retargeting step. We\ncompare to previous work [Fried et al . 2019; Thies et al . 2020] and\nto ground-truth video. We follow the study design used in previous\ntalking-head synthesis research [Fried et al . 2019; Kim et al . 2018].\nSpeci\ue004cally, participants see one video at a time in randomized order\nand are asked to rate the statement \u201cThis video clip looks real to\nme\u201d on a 5-point Likert scale ranging from strongly disagree (1) to\nstrongly agree (5). All videos used in our studies are available in the\nsupplemental material.\nUser study 1: Short Phrases (1 \u2013 4 words). Short phrases are the\nmain type of result shown in Fried et al. [2019]. Such edits are useful\nfor minor \ue004x-ups on existing sentences. In this study we compare\nour automatic synthesis results (\u201cOurs\u201d) to 3 versions of Fried et\nal. [2019]. (1) Their method with the same amount of target video as\nused by our tool, which is less than 5 minutes in all cases. (2) Their\nmethod with 1 hour of target video, which is their recommended\namount, and more than 12 times the amount we use. (3) A version of\ntheir method with our fast phoneme search and stitching algorithm\n(Section 3.3), but with 1 hour of target video (\u201cModi\ue004ed Fried et\nal. [2019]\u201d in Section 5.1), to evaluate the e\ue006ect of ablating our neural\nretargeting step. We also compare to ground truth video recordings.\nWe recruited 110 participants to view 25 videos each (5 conditions\nfor each of 5 edits). We report Likert scale responses in Table 2\n(\u201cShort Phrase\u201d). The di\ue006erences between all pairs, except \u201cOurs\u201d\nvs. \u201cModi\ue004ed Fried\u201d, are statistically signi\ue004cant. All p-values have\n, Vol. 1, No. 1, Article . Publication date: November 2020.\n12 \u2022 Yao, et al.\nbeen adjusted for multiple testing and are reported in supplemental\nmaterials.\nOur tool outperforms Fried et al. [2019] both when using 5 min-\nutes of data and 1 hour of data. We believe this is due to our results\nhaving more accurate mouth closures and better temporal coherency\nin mouth motions. Results are similar for our tool and Modi\ue004ed Fried,\nindicating that our neural retargeting step does not have much neg-\native e\ue006ect on result quality. Together these results also suggest that\nour fast phoneme search with stitching that forces closed mouths\nfor \\m, \\b, and \\p phonemes leads to higher-quality synthesis than\nthe slow phoneme search and stitching approach used originally\nby Fried et al. [2019]. Although a gap still remains between our\nsynthesized results and ground-truth videos, our results for short\nedits are rated as real almost two thirds of the time.\nUser study 2: Full Sentences (6 \u2013 9 words). Full sentence synthesis is\nmore challenging, since longer synthesis equates to a larger chance\nof inaccurate matches and synthesis artifacts. However, synthesiz-\ning full sentences as opposed to short phrases opens up more use\ncases (Section 4). Investigating full-sentence synthesis also empha-\nsizes the quality di\ue006erences between methods. The conditions in\nthis user study are the same as for user study 1. We recruited 153\nparticipants to view 25 videos each (5 conditions for each of 5 sen-\ntences). We report Likert scale responses in Table 2 (\u201cFull Sentence\u201d).\nThe di\ue006erences between all pairs, except \u201cFried < 5 min\u201d vs. \u201cFried\n\n    1 hr\u201d, are statistically signi\ue004cant. All p-values have been adjusted\n    for multiple testing and are reported in supplemental materials.\n    Our tool produces the highest-quality results, followed by Modi-\n    \ue004ed Fried with over 1 hour of data, then by Fried et al. [2019]. Similar\n    to user study 1, here our results have better mouth closures and\n    smoother mouth motions than Fried et al. [2019]. It is worth noting\n    that our results are even better than Modi\ue004ed Fried. We believe this\n    is because our tool has a higher-quality source repository which be-\n    comes more salient when the edits are long. It shows the advantage\n    of our approach to decouple source repository from target video,\n    as data quality improvements to the repository can bene\ue004t many\n    di\ue006erent target videos. The one-time cost of building a high-quality\n    repository amortizes across all the edits that use it.\n    User study 3: our tool vs Neural Voice Puppetry [Thies et al . 2020].\n    The third user study compares our results to those of Neural Voice\n    Puppetry [Thies et al . 2020], where we show viewers videos gen-\n    erated by the two methods from the same audio speech track. We\n    recruited 90 participants to view 8 videos each (4 from each of the\n    two methods, Table 3). The audio tracks used in the videos are not\n    the actor\u2019s real voice, and we believe this is the predominant reason\n    for overall lower scores (for both methods). The di\ue006erence between\n    conditions is not statistically signi\ue004cant, and our results have sim-\n    ilar mean scores to those of NVP. Nevertheless, as mentioned in\n    Section 5.3, closely examining the videos generated by the two ap-\n    proaches, we \ue004nd that our method often does a better job of closing\n    the mouth on \\m, \\b, and \\p phonemes. We also note that while our\n    user studies evaluate our automatic results, unlike NVP, our tool\n    also provides re\ue004nement and performance controls that can be used\n    to improve results over the course of an interactive editing session.\n    Likert response (%)\n    Condition Length 5 4 3 2 1 Mean \u2018Real\u2019\n    Short Phrase Fried et al. [2019] < 5 min 19.5 28.0 11.3 22.1 19.0 3.1 47.6%\n    Fried et al. [2019] > 1 hr 24.1 31.5 13.7 20.0 10.7 3.4 55.6%\n    Modi\ue004ed Fried > 1 hr 27.5 39.9 13.3 13.9 5.4 3.7 67.4%\n    Ours < 5 min 30.37 34.2 14.1 15.2 5.7 3.7 64.9%\n    Ground truth n/a 40.1 37.8 11.5 9.8 0.9 4.1 77.8%\n    Full Sentence Fried et al. [2019] < 5 min 14.7 22.9 11.1 25.0 26.3 2.7 37.6%\n    Fried et al. [2019] > 1 hr 14.6 22.5 12.2 26.3 24.5 2.8 37.1%\n    Modi\ue004ed Fried > 1 hr 16.5 31.5 14.1 20.7 17.2 3.1 48.0%\n    Ours < 5 min 17.9 38.2 16.0 20.1 7.7 3.4 56.2%\n    Ground truth n/a 39.0 42.1 7.3 8.8 2.9 4.1 81.1%\n    Table 2. Results from user studies on short phrases and full sentences.\n    The \u201cLength\u201d column shows the length of the input target video for each\n    method. We compare to previous work [Fried et al . 2019] and to ground-\n    truth recordings. We report percentage of each answer on a 5-Point Likert\n    scale, the mean score, and percent of videos that received a score of 4 or\n    5 (\u2018real\u2019). The di\ue003erence between conditions is significant in both studies\n    (Kruskal-Wallis test, \ue00c < 10 \u221220 each). A followup Tukey\u2019s range test shows\n    that all pairwise comparisons are statistically significant (\ue00c < 0 . 008 each)\n    except for \u201cOurs\u201d vs. \u201cModified Fried\u201d for short edits and \u201cFried < 5 min\u201d\n    vs. \u201cFried > 1 hr\u201d for full sentences. Note that Tukey\u2019s procedure adjusts\n    the p-values for multiple comparisons. We report all adjusted p-values in\n    the supplemental materials. Using our fast phoneme search and stitching\n    algorithm improves results from Fried et al. [2019]. Our tool outperforms\n    the method of Fried et al. [2019] when they use the same amount of target\n    video data, and when they use 12x the amount of data. For full-sentence\n    synthesis, our tool also outperforms Modified Fried even while they use 12x\n    the amount of data.\n    Likert response (%)\n    Condition Length 5 4 3 2 1 Mean \u2018Real\u2019\n    Ours < 5 min 6.0 9.6 10.6 26.2 47.5 2.0 15.6%\n    NVP < 5 min 6.0 10.6 11.3 24.8 47.4 2.0 16.6%\n    Table 3. Results from user study on our tool and Neural Voice Puppetry\n    (NVP) [Thies et al . 2020]. We compare to NVP and report percentage of each\n    answer on a 5-Point Likert scale, as well as mean score and percent of videos\n    that received a score of 4 or 5 (\u2018real\u2019). Our tool and NVP received similar\n    mean scores and the di\ue003erence is not statistically significant (Kruskal-Wallis\n    test, \ue00c = 0.84).\n    Automatic Metrics. For videos in user study 1 and 2 where we\n    have ground-truth recordings, we further evaluate the results using\n    automatic metrics against grouth truth videos. To measure recon-\n    struction quality, we compute structural similarity index (SSIM\n    [Zhou Wang et al . 2004]) and peak signal-to-noise ratio (PSNR). To\n    measure lip motion quality, we compute the Landmarks Distances\n    (LMD) [Chen et al . 2018]. We report the results in Table 4. Although\n    our methods achieve favorable scores on many of these metrics, the\n    score di\ue006erences are quite small and visual di\ue006erences can be subtle.\n    We believe the user studies provide a better and more trustworthy\n    measure of video quality.\n    , Vol. 1, No. 1, Article . Publication date: November 2020.\n    Iterative Text-based Editing of Talking-heads Using Neural Retargeting \u2022 13\n    Condition Length SSIM PSNR LMD\n    Short Phrase Fried et al. [2019] < 5 min 0.89929 25.08146 4.57492\n    Fried et al. [2019] > 1 hr 0.89925 25.09057 4.47108\n    Modi\ue004ed Fried > 1 hr 0.89934 25.09670 4.38139\n    Ours < 5 min 0.89921 25.09349 4.57060\n    Full Sentence Fried et al. [2019] < 5 min 0.96339 32.74630 3.75896\n    Fried et al. [2019] > 1 hr 0.96552 32.94852 3.67367\n    Modi\ue004ed Fried > 1 hr 0.97578 35.00073 2.90959\n    Ours < 5 min 0.97630 35.12082 3.18670\n    Table 4. Results of automatic metrics. We compare our results to 3 versions\n    of Fried et al. [2019] for both short phrase edits and full sentence syntheses.\n    Best score is bolded. Our tool tops SSIM and PSNR for full sentence syn-\n    theses and ranks second a\ue004er Modified Fried for LMD on full sentence and\n    PSNR on short edits.\n    6 LIMITATIONS AND FUTURE WORK\n    We have demonstrated an iterative text-based tool for editing talking-\n    head dialogue and performance that can be applied to many real-\n    world editing scenarios in which only a few minutes of target actor\n    video is available. However, our approach does have several limita-\n    tions that could be addressed in future work.\n    Further reduce feedback loop time. Our tool currently requires\n    about 30 seconds to synthesize a typical 5 word edit. While this feed-\n    back loop time allows users to try a variety of edits and re\ue004nements,\n    seeing a synthesized result immediately (in real-time) would allow\n    even more iteration and exploration of design space. As noted in\n    Section 3.6, parallelization of our fast phoneme search and neural\n    rendering steps as well as streaming playback of the synthesized\n    video could reduce the feedback loop time signi\ue004cantly.\n    Improve quality of synthesis results. Although our method com-\n    pares favorably in quality with previous talking-head synthesis\n    techniques (Section 5.4), there is still a gap in realism between our\n    results and ground-truth videos. As our method relies on a rich\n    repository of source video to provide mouth motions for phoneme\n    coarticulations, it may be possible to improve synthesis by develop-\n    ing higher-quality repositories. One approach may be to leverage\n    existing work in text-driven 3D human mouth animation [Edwards\n    et al . 2016] to render unlimited amounts of mouth motions to serve\n    as the repository. Another direction is to build multiple reposito-\n    ries of many di\ue006erent source actors and then given a target video,\n    develop techniques to pick the best source actor for the target.\n    Performance controls over full face. Our current approach focuses\n    on synthesizing lip motions that match the target edit. While our\n    tool o\ue006ers controls for inserting mouth gestures and changing the\n    speaking style, the e\ue006ects of these controls are limited to the lower\n    part of the face. Others have demonstrated techniques for controlling\n    more of the head, including the ability to change head pose, gaze\n    direction and whole face expressions [Kim et al . 2018] However,\n    these techniques often introduce artifacts in the hair and with the\n    clothes at the neckline. Adding such full face controls in an artifact-\n    free manner remains an open research direction.\n    Previsualizing dialogue using existing \ue001lm scenes. When writing\n    dialogue, scriptwriters have to imagine the sound and appearance\n    of the scene. Using our video editing tool with a catalog of video\n    from existing \ue004lm scenes might allow such scriptwriters to quickly\n    visualize the dialogue in di\ue006erent settings and with di\ue006erent actors.\n    Users might search for scenes based on their settings and actors\n    using a tool like SceneSkim [Pavel et al . 2015] and our tool could be\n    used to insert the new dialogue.\n    7 ETHICAL CONSIDERATIONS\n    Our editing tool is designed to enable an iterative work\ue007ow for\n    removing \ue004ller words, adjusting phrasing, or correcting mistakes in\n    a talking-head video. While such tools can facilitate content creation\n    and storytelling, tools like ours, that let users manipulate what a\n    target actor is saying, can also be misused. We follow the guidelines\n    suggested by Fried et al. [2019] for ethically using such tools. (1)\n    Video generated by our tool should be transparent about the fact\n    that it has been manipulated. (2) Actors must give consent to any\n    manipulation before a resulting video is shared widely.\n    We also recognize that these guidelines alone will not stop bad\n    actors from using tools like ours to create false statements and slan-\n    der others. Therefore, it is also critical for researchers to continue\n    developing tools for detecting, \ue004ngerprinting, and verifying such\n    video manipulation. Openly publishing the technical details of our\n    tool can increase public awareness and help detection e\ue006orts. Ul-\n    timately these issues may also require regulations and laws that\n    penalize misuse while allowing creative and consensual use cases.\n    8 CONCLUSION\n    Iterative editing is central to many content-creation tasks and is\n    especially common in video editing. We have shown how to enable\n    such iterative editing in the context of editing talking-head video\n    using a text-based interface that allows changes to wording and\n    facial performance while providing re\ue004nement controls. Whether\n    an editor trying di\ue006erent ways to phrase the dialogue in a \ue004lm,\n    developing dialogue for a conversational agent, or correcting a\n    mistake in a lecture, such iteration is often essential for \ue004nding the\n    most appropriate result. We believe such tools that facilitate video\n    editing can democratize content-creation and enable many more\n    people to tell their stories.\n    REFERENCES\n    2020a. Google Cloud Speech to Text API. https://cloud.google.com/speech-to-text\n    2020b. Google Cloud Text to Speech API. https://cloud.google.com/text-to-speech\n    2020. Lyrebird AI. https://www.descript.com/lyrebird-ai\n    2020. Rev. https://rev.com\n    Mart\u00edn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig\n    Citro, Greg S. Corrado, Andy Davis, Je\ue006rey Dean, Matthieu Devin, Sanjay Ghe-\n    mawat, Ian Goodfellow, Andrew Harp, Geo\ue006rey Irving, Michael Isard, Yangqing\n    Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dande-\n    lion Man\u00e9, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster,\n    Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent\n    Vanhoucke, Vijay Vasudevan, Fernanda Vi\u00e9gas, Oriol Vinyals, Pete Warden, Martin\n    Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015. TensorFlow: Large-\n    Scale Machine Learning on Heterogeneous Systems. https://www.tensor\ue007ow.org/\n    Software available from tensor\ue007ow.org.\n    Shruti Agarwal, Hany Farid, Ohad Fried, and Maneesh Agrawala. 2020. Detecting Deep-\n    Fake Videos from Phoneme-Viseme Mismatches. In Workshop on Media Forensics at\n    CVPR. Seattle, WA, USA.\n    , Vol. 1, No. 1, Article . Publication date: November 2020.\n    14 \u2022 Yao, et al.\n    Hadar Averbuch-Elor, Daniel Cohen-Or, Johannes Kopf, and Michael F. Cohen. 2017.\n    Bringing Portraits to Life. 36, 6 (Nov. 2017), 196:1\u201313. https://doi.org/10.1145/\n    3130800.3130818\n    Floraine Berthouzoz, Wilmot Li, and Maneesh Agrawala. 2012. Tools for Placing Cuts\n    and Transitions in Interview Video. ACM Trans. Graph. 31, 4, Article 67 (July 2012),\n    8 pages. https://doi.org/10.1145/2185520.2185563\n    Volker Blanz and Thomas Vetter. 1999. A Morphable Model for the Synthesis of\n    3D Faces. In Proceedings of the 26th Annual Conference on Computer Graphics and\n    Interactive Techniques (SIGGRAPH \u201999). ACM Press/Addison-Wesley Publishing Co.,\n    USA, 187\u2013194. https://doi.org/10.1145/311535.311556\n    Christoph Bregler, Michele Covell, and Malcolm Slaney. 1997. Video Rewrite: Driving\n    Visual Speech with Audio. In Proceedings of the 24th Annual Conference on Computer\n    Graphics and Interactive Techniques (SIGGRAPH \u201997). ACM Press/Addison-Wesley\n    Publishing Co., USA, 353\u2013360. https://doi.org/10.1145/258734.258880\n    Yao-Jen Chang and Tony Ezzat. 2005. Transferable Videorealistic Speech Animation.\n    143\u2013151. https://doi.org/10.1145/1073368.1073388\n    Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2018. Lip\n    movements generation at a glance. In Proceedings of the European Conference on\n    Computer Vision (ECCV). 520\u2013535.\n    Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2019. Hierarchical cross-\n    modal talking face generation with dynamic pixel-wise loss. In Proceedings of the\n    IEEE Conference on Computer Vision and Pattern Recognition. 7832\u20137841.\n    Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017. You said that?. In\n    British Machine Vision Conference.\n    Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016. JALI: an animator-\n    centric viseme model for expressive lip synchronization. ACM Transactions on\n    Graphics (TOG) 35, 4 (2016), 1\u201311.\n    Tony Ezzat, Gadi Geiger, and Tomaso Poggio. 2002. Trainable Videorealistic Speech\n    Animation. 21, 3 (July 2002), 388\u2013398. https://doi.org/10.1145/566654.566594\n    Ohad Fried and Maneesh Agrawala. 2019. Puppet Dubbing. In Eurographics Symposium\n    on Rendering - DL-only and Industry Track, Tamy Boubekeur and Pradeep Sen (Eds.).\n    The Eurographics Association. https://doi.org/10.2312/sr.20191220\n    Ohad Fried, Ayush Tewari, Michael Zollh\u00f6fer, Adam Finkelstein, Eli Shechtman, Dan B\n    Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019.\n    Text-Based Editing of Talking-Head Video. ACM Trans. Graph. 38, 4, Article 68 (July\n    2019), 14 pages. https://doi.org/10.1145/3306346.3323028\n    John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, and David S Pallett.\n    1993. DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST\n    speech disc 1-1.1. NASA STI/Recon technical report n 93 (1993).\n    Pablo Garrido, Levi Valgaerts, Ole Rehmsen, Thorsten Thormaehlen, Patrick P\u00e9rez,\n    and Christian Theobalt. 2014. Automatic Face Reenactment. 4217\u20134224. https:\n    //doi.org/10.1109/CVPR.2014.537\n    Pablo Garrido, Levi Valgaerts, Hamid Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick\n    P\u00e9rez, and Christian Theobalt. 2015. VDub: Modifying Face Video of Actors for\n    Plausible Visual Alignment to a Dubbed Audio Track. 34, 2 (May 2015), 193\u2013204.\n    https://doi.org/10.1111/cgf.12552\n    Pablo Garrido, Michael Zollh\u00f6fer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick\n    P\u00e9rez, and Christian Theobalt. 2016. Reconstruction of Personalized 3D Face Rigs\n    from Monocular Video. ACM Trans. Graph. 35, 3, Article 28 (May 2016), 15 pages.\n    https://doi.org/10.1145/2890493\n    Jiahao Geng, Tianjia Shao, Youyi Zheng, Yanlin Weng, and Kun Zhou. 2018. Warp-\n    guided GANs for Single-photo Facial Animation. In SIGGRAPH Asia 2018 Technical\n    Papers (Tokyo, Japan) (SIGGRAPH Asia \u201918). ACM, New York, NY, USA, Article 231,\n    231:1\u2013231:12 pages. http://doi.acm.org/10.1145/3272127.3275043\n    Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen,\n    Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al . 2018. Transfer learning\n    from speaker veri\ue004cation to multispeaker text-to-speech synthesis. In Advances in\n    neural information processing systems. 4480\u20134490.\n    Ira Kemelmacher-Shlizerman, Aditya Sankar, Eli Shechtman, and Steven M. Seitz. 2010.\n    Being John Malkovich. 341\u2013353. https://doi.org/10.1007/978-3-642-15549-9_25\n    Hyeongwoo Kim, Mohamed Elgharib, Michael Zollh\u00f6fer, Hans-Peter Seidel, Thabo\n    Beeler, Christian Richardt, and Christian Theobalt. 2019. Neural Style-Preserving\n    Visual Dubbing. ACM Trans. Graph. 38, 6, Article 178 (Nov. 2019), 13 pages. https:\n    //doi.org/10.1145/3355089.3356500\n    H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, N. Nie\u00fbner, P. P\u00e9rez, C. Richardt, M.\n    Zollh\u00f6fer, and C. Theobalt. 2018. Deep Video Portraits. ACM Transactions on\n    Graphics 2018 (TOG) (2018).\n    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization.\n    In 3rd International Conference on Learning Representations, ICLR 2015, San Diego,\n    CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann\n    LeCun (Eds.). http://arxiv.org/abs/1412.6980\n    Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh,\n    Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron Courville. 2019.\n    MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis.\n    arXiv:1910.06711 [eess.AS]\n    Rithesh Kumar, Jose Sotelo, Kundan Kumar, Alexandre de Brebisson, and Yoshua Bengio.\n    2017. ObamaNet: Photo-realistic lip-sync from text. arXiv:1801.01442 [cs.CV]\n    VI Levenshtein. 1966. Binary Codes Capable of Correcting Deletions, Insertions and\n    Reversals. In Soviet Physics Doklady, Vol. 10. 707.\n    Kang Liu and Joern Ostermann. 2011. Realistic facial expression synthesis for an\n    image-based talking head. https://doi.org/10.1109/ICME.2011.6011835\n    Wesley Mattheyses, Lukas Latacz, and Werner Verhelst. 2010. Optimized photorealistic\n    audiovisual speech synthesis using active appearance modeling. In Auditory-Visual\n    Speech Processing. 8\u20131.\n    Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the di\ue005culty of training\n    recurrent neural networks. In International conference on machine learning. 1310\u2013\n    1318.\n    Amy Pavel, Dan B Goldman, Bj\u00f6rn Hartmann, and Maneesh Agrawala. 2015. Sce-\n    neskim: Searching and browsing movies using synchronized captions, scripts and\n    plot summaries. In Proceedings of the 28th Annual ACM Symposium on User Interface\n    Software & Technology. ACM, 181\u2013190.\n    A. Pumarola, A. Agudo, A.M. Martinez, A. Sanfeliu, and F. Moreno-Noguer. 2019.\n    GANimation: One-Shot Anatomically Consistent Facial Animation. (2019).\n    Steve Rubin, Floraine Berthouzoz, Gautham J Mysore, Wilmot Li, and Maneesh\n    Agrawala. 2013. Content-based tools for editing audio stories. In Proceedings of\n    the 26th annual ACM symposium on User interface software and technology. ACM,\n    113\u2013122.\n    Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi. 2019. Talking\n    Face Generation by Conditional Recurrent Adversarial Network. In Proceedings of\n    the Twenty-Eighth International Joint Conference on Arti\ue001cial Intelligence, IJCAI-19.\n    International Joint Conferences on Arti\ue004cial Intelligence Organization, 919\u2013925.\n    https://doi.org/10.24963/ijcai.2019/129\n    Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. 2017.\n    Synthesizing Obama: Learning Lip Sync from Audio. ACM Trans. Graph. 36, 4,\n    Article 95 (July 2017), 13 pages. https://doi.org/10.1145/3072959.3073640\n    Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan\n    Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nie\u00fbner,\n    Rohit Pandey, Sean Fanello, Gordon Wetzstein, Jun-Yan Zhu, Christian Theobalt,\n    Maneesh Agrawala, Eli Shechtman, Dan B. Goldman, and Michael Zollh\u00f6fer. 2020.\n    State of the Art on Neural Rendering. Computer Graphics Forum (2020). https:\n    //doi.org/10.1111/cgf.14022\n    Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias\n    Nie\u00fbner. 2020. Neural Voice Puppetry: Audio-driven Facial Reenactment. In European\n    Conference on Computer Vision. Springer, 716\u2013731.\n    Justus Thies, Michael Zollh\u00f6fer, Marc Stamminger, Christian Theobalt, and Matthias\n    Nie\u00fbner. 2016. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos.\n    2387\u20132395. https://doi.org/10.1109/CVPR.2016.262\n    A\u00e4ron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals,\n    Alexander Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016.\n    WaveNet: A Generative Model for Raw Audio. In Arxiv. https://arxiv.org/abs/1609.\n    03499\n    Daniel Vlasic, Matthew Brand, Hanspeter P\ue004ster, and Jovan Popovi\u0107. 2005. Face Transfer\n    with Multilinear Models. 24, 3 (July 2005), 426\u2013433. https://doi.org/10.1145/1073204.\n    1073209\n    Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2018. End-to-End Speech-\n    Driven Facial Animation with Temporal GANs. arXiv:1805.09313 [eess.AS]\n    Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2019. Realistic Speech-\n    Driven Facial Animation with GANs. International Journal of Computer Vision (13\n    Oct 2019). https://doi.org/10.1007/s11263-019-01251-8\n    Robert A Wagner and Michael J Fischer. 1974. The string-to-string correction problem.\n    Journal of the ACM (JACM) 21, 1 (1974), 168\u2013173.\n    Lijuan Wang, Wei Han, Frank Soong, and Qiang Huo. 2011. Text-driven 3D Photo-\n    Realistic Talking Head. In INTERSPEECH 2011 (interspeech 2011 ed.). International\n    Speech Communication Association. https://www.microsoft.com/en-us/research/\n    publication/text-driven-3d-photo-realistic-talking-head/\n    O. Wiles, A.S. Koepke, and A. Zisserman. 2018. X2Face: A network for controlling\n    face generation by using images, audio, and pose codes. In European Conference on\n    Computer Vision.\n    Jiahong Yuan and Mark Liberman. 2008. Speaker identi\ue004cation on the SCOTUS corpus.\n    In In Proceedings of Acoustics 2008. Citeseer.\n    Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019.\n    Few-Shot Adversarial Learning of Realistic Neural Talking Head Models.\n    arXiv:1905.08233 [cs.CV]\n    Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019. Talking Face\n    Generation by Adversarially Disentangled Audio-Visual Representation. In AAAI\n    Conference on Arti\ue001cial Intelligence (AAAI).\n    Zhou Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. 2004. Image quality\n    assessment: from error visibility to structural similarity. IEEE Transactions on Image\n    Processing 13, 4 (2004), 600\u2013612.\n    , Vol. 1, No. 1, Article . Publication date: November 2020.\n\n"},{"assistant_initiates":false,"asst_text":"The 3D head model registration process involves fitting a parametric model that captures the facial structure and expressions of the actor's face in each frame of a video. Here is a detailed breakdown of the steps and components involved in this process:\n\n1. **Parametric 3D Head Model**: This model is typically based on a pre-existing framework like the Basel Face Model or a similar morphable model. It includes several parameters that together define the shape, texture, and expression of the face.\n\n2. **Face Tracking and Initialization**: The process begins by detecting facial landmarks in the video frame, which are then used to initialize the 3D head model. The landmarks help in aligning the model with the face in the image, providing a rough fit which will be refined.\n\n3. **Model Fitting**: The fitting process involves optimizing the parameters of the 3D head model to minimize the difference between the projected 3D model and the observed 2D image. This typically involves:\n    - **3D Geometry Parameters**: These define the overall shape of the face and are usually kept constant over a sequence of frames to maintain identity consistency.\n    - **Reflectance Parameters**: These define the skin tone and texture. Like geometry parameters, they are typically constant throughout the video for a given actor.\n    - **Expression Parameters**: These parameters vary per frame and capture facial expressions, especially those of the lips and eyes.\n    - **Pose Parameters**: Define the head's orientation in the frame and change dynamically as the head moves.\n    - **Illumination Parameters**: Capture the lighting conditions affecting the face. These are also variable per frame to account for changes in lighting as the face moves.\n\n4. **Optimization Process**: Fitting is usually achieved through optimization algorithms that adjust the model parameters to best match the facial features in the video frame. This involves minimizing a loss function, typically accounting for the difference in appearance between the synthesized and real image, as well as constraints to ensure realistic outputs (e.g., smoothness, consistency with known expression constraints).\n\n5. **Temporal Consistency**: The fitting process includes ensuring temporal consistency so that the movements of the model appear smooth over time. This may involve using temporal filters or constraints that ensure that rapid, unrealistic changes between frames are minimized.\n\n6. **Output Parameters**: Once the model is fitted, it provides a set of parameters for each frame of the video, which can be used to reconstruct the face in 3D space. These parameters allow manipulation and retargeting of expressions from one actor to another while maintaining the identity of the target.\n\nIn this context, the fitted parameters are crucial for retargeting facial expressions from a source actor to a target actor during the synthesis phase of video editing, enabling precise alignment of lip-sync movements and expressions.","behavior":"degenerate_repetition","conversation_hash":"f02091fd279e2bca5db1ca1e03b806f3","conversation_redacted":false,"conversation_toxic":false,"country":"Sweden","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"English","model":"gpt-4o-2024-08-06","n_chars":3045,"n_turns_in_conversation":13,"n_words":463,"source_row_offset":29634,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:49:18 GMT","turn_id":"6c2e4a53576172e6bd8fe515","turn_index":1,"user_text":"explain 3d head model registration in detail"},{"assistant_initiates":false,"asst_text":"Yes, the explanation provided aligns with the methodology typically involved in 3D head model registration as it would relate to the goals outlined in the paper \"Iterative Text-based Editing of Talking-heads Using Neural Retargeting.\" The paper uses a parametric 3D head model to capture the facial expressions of an actor and enable text-based video editing by transferring these expressions to the target video. This involves:\n\n- Fitting a 3D morphable model to each frame of video to derive parameters for geometry, reflectance, expression, pose, and illumination.\n- Using a monocular head tracker to automate this fitting process, which registers the model to the video frames, capturing the necessary dynamics of facial movements.\n- Ensuring that certain parameters (like expression and pose) are allowed to vary over time, whereas others (like geometry and reflectance) remain consistent for an actor across the video.\n\nThis registration process outputs a sequence of parameters that can be used in the synthesis pipeline for retargeting expressions between actors, a critical step outlined in the paper. Hence, the detailed explanation reflects the process as described and implied in the paper's methodology.","behavior":"degenerate_repetition","conversation_hash":"f02091fd279e2bca5db1ca1e03b806f3","conversation_redacted":false,"conversation_toxic":false,"country":"Sweden","empty_output":false,"evidence_codes":[],"label":"absent","label_present":false,"label_qualified":false,"label_source":"Final dense release","language":"English","model":"gpt-4o-2024-08-06","n_chars":1253,"n_turns_in_conversation":13,"n_words":191,"source_row_offset":29635,"source_shard":"turns-00039.parquet","timestamp":"Fri, 25 Oct 2024 14:49:18 GMT","turn_id":"3f07eca9db47939280b6d6c4","turn_index":2,"user_text":"does you answer align with the paper"}],"shard":{"index":39,"name":"turns-00039.parquet","rows":41893}}
