From bdd6a09239a3ff4e7a15f12c1d8b3d93394d7efa Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 16 Mar 2023 21:08:54 -0500 Subject: [PATCH 01/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 94 +++++++++++++++++++ 1 file changed, 94 insertions(+) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 6ff6958..2671937 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -7,6 +7,7 @@ parent: Technical Committee --- # Notes of the Technical Committee Meeting - March 10, 2023 + **The Meeting Began at 10:07am EST.** **Attendees: J. Stine, N. Southern, B. Epstein, O. Coleman, D. Dahl, J. Larson, T. Martens** @@ -58,6 +59,99 @@ parent: Technical Committee * Data Packets * Context and History +*Mr. Stine noted that the Information Systems Authority (ISA) of Estonia, creators of the Bürokratt virtual assistant, have agreed to be one of the external partners whom Dr. Dahl identified above.* + +*Mr. Martens noted that the survey will be used to develop a tool that can be used when friends and partners of OVON are encountered - a way to involve them in prototypes. The survey/interview guide will facilitate conversations. It is a peer-reviewed effort.* + +*Mr. Stine added a mention of Ms. Olga Howard's approach to use cases; this will continue to be valuable. The team will also continue having discussions with industry people on the outside. However, Ms. Howard's work is instrumental in connecting external use cases to the work led by the interoperability team.* + +*Mr. Martens noted that the Estonian government represents a key piece of the equation for Open Voice regarding the establishment of a government partner; what remains is for OVON to find multiple technology providers - given the discussions about generative AI, Microsoft AI is one possibility, Nuance is another. OVON should also have corporate users of Voice Assistant technologies such as Schwarz Grüppe or Volkswagen.* + +*Dr. Dahl then reviewed the Interoperability Roadmap, as set up in Trello. and presented a slide with the group's partner activities, main activity and the future demos and draft specifications. She then presented a timeline view of the previous slide, and noted that the group is nearly finished with Demo 1 and close to finishing the Dialog Packet draft. Also forthcoming is Demo 2, which has broader capability than Demo 1. Demo 1 is two independent agents delegating tasks back and forth to one another, with a simple use case and speech interaction (a "hello world" demo). Demo 2 - the main addition is use of more specifications, such as negotiation and cooperation between two agents. Demo 3 will introduce channeling pattern where the second agent - the one delegated to, that does the actual work - asks the first agent to speak and listen for it. The agent that is being delegated to is conducting the dialog, but using the agent to which the user is speaking as its voice. By the end of Demo 3 the group will be able to demo two patterns.* + + +## Headline Deliverables 1H 2023 - J. Stine ## + +*Mr. Stine presented a slide detailing OVON's deliverables for the first half of 2023:* + +* Build Qualified Sponsor-Contributor Participant Pipeline + * 1H Success Metric: Second-Level Conversations with >10 potential >$10k + contributors/sponsors + * A new category of support for OVON: a 1x or 1Y contribution to OVON without + being a member of the Linux Foundation. This is a way for smaller firms to support + the Open Voice Network. This will open up additional revenue streams. +* Launch TrustMark Program - launched by Ms. Coleman on March 29, 2023 in a webinar +* Publish-Demo First Interoperability Specifications. + +*Mr. Stine also noted that the group is on track to deliver. Deliverables 2 and 3 above make Deliverable 1 doable as they build organizational credibility.* + +## Ecosystem for Participation, Development 1H 2023 - J. Stine ## + +*Mr. Stine discussed potential vertical prospects for OVON, in areas including Automotive, Healthcare, Social Gaming, Financial Services, Retail, Restaurant Hospitality, Government and Smart Home. He noted that the same issues face all, but they differ in use cases and metrics. He also referenced several horizontal value propositions that cut across industries, including Customer Service, Enterprise Content Management and Standardized Transaction/Authentication. In the latter area, there is not a standard method for voice-based authentication and transaction which is a concern in all the consumer-facing industries. Mr. Stine also noted the presence of a development ecosystem, and different small providers of solutions in various industries, clustered. Different firms are also doing strategy and solution, security/authentication, open voice platforms, language models and big monster tech with specific interests in conversational AI.* + +*Mr. Stine then outlined the distinction between OVON members, sponsors and contributors, each class an integral part of the Open Voice ecosystem.* + +*Mr. Stine stressed the criticality at this time of OVON transitioning from reactive to proactive as it shifts to outreach to individuals in all of the said categories. Progress on this outreach will be reported to the technical committee.* + +## Authentication Study Group Report - B. Epstein ## + +*The Authentication Study Group launched on Fri., Mar. 3, 2023, output posted on slack and the google drive, with the following guidelines:* + +* Objective: Identify and document the technical requirements for identifying and authenticating voice assistants. +* Two primary questions: + * How do we identify? + * How do we authenticate? +* Areas not covered - that may need additional study groups: + * Discovery + * Disambiguation (similar sounding names - we are not concerned with this yet) + * Arbitration (how a voice system chooses which assistant it connects to) + * User identification and authentication + * Goal is to have a first draft of the specifications document ready by the end of April. + +*The group is in a 'call for comments' phase - to collect ideas for requirements about the identification and authentication of voice assistants.* + +*Mr. Epstein asked if anyone is aware of other initiatives working on authentication of voice assistants or of users to please let him know.* + +*Mr. Martens noted the many ongoing discussions about Large Language Models and how they should identify themselves - about training data, biases, and the like. For that reason, this exercise can benefit many stakeholders.* + + +## Trust Mark Initiative Updates - O. Coleman + +* On deck to have program launch on 3/29/23 +* Helpful discussion regarding branding Trust Mark with the legal team at Linux Foundation about the need for copyright or lack thereof +* The current working name is 'The Trust Mark Initiative of the Open Voice Network.' +* A webinar is scheduled for March 29th. The website will be updated and social media outreach is planned. +* Also on the 29th - the first module of the LF training course will be previewed with a teaser. +* Letters of invitation for a board of advisors are going out - a group of people in the industry who can help with the program, providing support and introduction to companies and entities that will offer support of the program +* Project Voice 2023 will include a formal announcement and a conversation with the voice AI leadership council. +* Excellent meeting with Deutsch Telekom. + * DT provided a view into its self-assessment tool, modeled in turn on the Trustworthy AI Framework from the European Union. + * Leveraging the work DT has done will get the Trust Mark initiative 90% of the way. + * The remainder of the work is the front matter of text changes, any training the team needs (for example, DT's educational training), etc. + * This may be a key investment area - an area where it makes sense to have a developer. It means using web forms but has IT support on the back end for authentication and data storage. + * It may mean putting a prototype on the website. + +* The group is aligning its code of ethics (to be announced - in draft mode now) - which covers the following principles + * Transparency + * Inclusivity + * Accountability + * Sustainability + * Privacy + * Compliance + +*Ms. Coleman also presented and called for a review of the code of ethics within the meeting. + +*At Project Voice, Bradley Metrock will be convening the Conversational A.I. Leadership Council. There will be 300 enterprise decision makers for this gathering. Those present will be asked by Metrock to sign the aforementioned OVON code of ethics. + +*OVON will be speaking at the Conversational A.I. Leadership Council and speaking to the principles it espouses in the said ethical statment. + +## Action Items - N. Southern: + +* Group - if anyone has ideas about requirements for identification of voice assistants, submit to Jon +* Group - if anyone is aware of other initiatives working on authentication of voice assistants or of users to please let Bruce know.* + + * Context and History + # Notes of the Open Voice Network Technical Committee Meeting - February 10, 2023 ### Attendees: J. Stine, N. Southern, O. Coleman, J. Larson, T. Martens, D. Dahl, E. Sewell, C. Wüttke, K. Brix, B. Epstein, N. Myers, E. Banzhaf. From a2489c895eb9229b0d7ce60b7c71b136a18cb12d Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 16 Mar 2023 21:12:31 -0500 Subject: [PATCH 02/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 12 +++++------- 1 file changed, 5 insertions(+), 7 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 2671937..3feb79d 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -5,7 +5,6 @@ layout: default title: Meeting Notes parent: Technical Committee --- - # Notes of the Technical Committee Meeting - March 10, 2023 **The Meeting Began at 10:07am EST.** @@ -139,18 +138,17 @@ parent: Technical Committee * Privacy * Compliance -*Ms. Coleman also presented and called for a review of the code of ethics within the meeting. +*Ms. Coleman also presented and called for a review of the code of ethics within the meeting.* -*At Project Voice, Bradley Metrock will be convening the Conversational A.I. Leadership Council. There will be 300 enterprise decision makers for this gathering. Those present will be asked by Metrock to sign the aforementioned OVON code of ethics. +*At Project Voice, Bradley Metrock will be convening the Conversational A.I. Leadership Council. There will be 300 enterprise decision makers for this gathering. Those present will be asked by Metrock to sign the aforementioned OVON code of ethics.* -*OVON will be speaking at the Conversational A.I. Leadership Council and speaking to the principles it espouses in the said ethical statment. +*OVON will be speaking at the Conversational A.I. Leadership Council and speaking to the principles it espouses in the said ethical statement.* ## Action Items - N. Southern: -* Group - if anyone has ideas about requirements for identification of voice assistants, submit to Jon -* Group - if anyone is aware of other initiatives working on authentication of voice assistants or of users to please let Bruce know.* +* Group - if anyone has ideas about requirements for identification of voice assistants, submit to Jon. +* Group - if anyone is aware of other initiatives working on authentication of voice assistants or of users to please let Bruce know. - * Context and History # Notes of the Open Voice Network Technical Committee Meeting - February 10, 2023 From 3902a47b8966fb73dbb611bcf16864fc046fa521 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 16 Mar 2023 22:12:01 -0500 Subject: [PATCH 03/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 30 +++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 3feb79d..b6e4678 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -144,10 +144,40 @@ parent: Technical Committee *OVON will be speaking at the Conversational A.I. Leadership Council and speaking to the principles it espouses in the said ethical statement.* +*Mr. Stine then noted that he finished a meeting with a German conversational AI development team in the midst of working on 'Charismatic Conversational AI' - with an emphasis on consumer-facing contexts such as call center and retail, hospitality and financial services. They feel OVON's work in ethical use is critical for them to espouse/support their ethical principles and to be able to say that said principles bely their work.* + +## Large Language Models and Generative A.I. - J. Stine + +*Mr. Stine shared a slide previously shown to the OVON Steering Committee on Feb. 23rd, about GAI and LLMs, and the tangible impact of these areas on the Open Voice Network. It covered the following questions and answers:* + +* Does it change OVON's mission? No, although it does increase industry-enterprise ethical use concerns and will, in time, increase perceived need for system-assistant interoperability. + +* Does it change OVON's work? Slightly. -Ethical use: means new use cases, larger harms. Interoperability: no, if language models have NLP-NKG front ends. Yes, if LM’s decoupled + +* Does it improve OVON's prospects? Yes, to a degree. Our leadership in ethical use even more important. It will, in time, increase perceived need forsystem-assistant interoperability. + +* Does it change OVON's messaging? Yes, to a degree. It’s the topic du jour we must reflect it in our white papers, presentations. + +*In response to this, Dr. Dahl suggested that there could be a widespread public perception that with the advent of ChatGPT there is no longer a need for interoperability, nor a development of natural language-based systems. She asked how to address this without getting into complicated waters. For instance - enterprise information that empowers enterprise agents and call centers and is hidden behind paywalls. Dr. Dahl then asked how we address this.* + +*Mr. Stine responded that mainstream media sources are drawing people to that conclusion, but that brand managers will not want their brands represented by OpenAI's large language model. To this end, Sean King of Veritone has just developed with his team a Generative AI process for creating domain-specific language models for brands - this is essential because a brand owner will want to manage the "truth" of that brand. In this new world, the process of obtaining information may move from queries to prompts. Every brand will need to create a domain-specific language model.* + +*Mr. Larson: two reasons why it is critical to do a better job on describing why we need interoperability:* +1. People will misperceive Generative AI's elimination of the need for interoperability +2. Some feel we already have PC-based interoperability. + +*Mr. Martens agreed with Dr. Dahl regarding the paradigm change about how important APIs are and how they are being addressed - and linked between systems - which becomes easier with LLMs. However, the current discussion sees all companies working to integrate the openAI APIs into their production systems.* + +*Mr. Epstein: if everything is being done with APIs, they do not need Open Voice. OVON is the Open Voice Network, not the Open API Network. OVON exists because humans will still want to interact with the system components through voice. Mr. Epstein re-stressed the criticality of finding a way to get voice assistants to interoperate.* + ## Action Items - N. Southern: +* Interoperability Team: develop an invitation message for webinar * Group - if anyone has ideas about requirements for identification of voice assistants, submit to Jon. * Group - if anyone is aware of other initiatives working on authentication of voice assistants or of users to please let Bruce know. +* Jon Stine - write an OVON blog statement on Generative AI and Large Language models. + +### Adjournment - Mr. Martens adjourned the meeting at 11:00am. # Notes of the Open Voice Network Technical Committee Meeting - February 10, 2023 From 0dd3957007719bfd8581905294fb97ad73c2eb0e Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 16 Mar 2023 22:12:52 -0500 Subject: [PATCH 04/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index b6e4678..0619bb2 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -27,7 +27,7 @@ parent: Technical Committee ## Planning of the Open Voice Interoperability Webinar - J. Larson, D. Dahl ## -*Dr. Dahl noted that the Interoperability Webinar is scheduled for Wed., Mar. 22nd, at 17:00h CET, 12:00pm EDT, and 9:00pm PDT. This will mainly cover the new interoperability standards on which Open Voice is working, especially the dialog packets. It will also cover the big picture of interoperability architecture with three corresponding patterns, and include the group's first demo of two independent interoperating agents sending delegation messages back and forth. Then a discussion will take place.* +*Dr. Dahl noted that the Interoperability Webinar is scheduled for Wed., Mar. 22nd, at 17:00h CET, 12:00pm EDT, and 9:00pm PDT. This will mainly cover the new interoperability standards on which Open Voice is working, especially the dialog packets. It will also cover the big picture of interoperability architecture with three corresponding patterns, and include the group's first demo of two independent interoperating agents sending delegation messages back and forth. Then a discussion will take place.* *Mr. Martens added that it will be an exciting opportunity if we can have a demo - to identify from the webinar a couple of smaller clips that we can share via social media.* From c5370baa82e6a2bcf5268488b0328fa3e03be1c2 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 16 Mar 2023 22:18:39 -0500 Subject: [PATCH 05/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 0619bb2..40117ec 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -27,7 +27,7 @@ parent: Technical Committee ## Planning of the Open Voice Interoperability Webinar - J. Larson, D. Dahl ## -*Dr. Dahl noted that the Interoperability Webinar is scheduled for Wed., Mar. 22nd, at 17:00h CET, 12:00pm EDT, and 9:00pm PDT. This will mainly cover the new interoperability standards on which Open Voice is working, especially the dialog packets. It will also cover the big picture of interoperability architecture with three corresponding patterns, and include the group's first demo of two independent interoperating agents sending delegation messages back and forth. Then a discussion will take place.* +*Dr. Dahl noted that the Interoperability Webinar is scheduled for Wed., Mar. 22nd, at 17:00h CET, 12:00pm EDT, and 9:00pm PDT. This will mainly cover the new interoperability standards on which Open Voice is working, especially the dialog packets. It will also cover the big picture of interoperability architecture with three corresponding patterns, and then include the group's first demo of two independent interoperating agents sending delegation messages back and forth. Then a discussion will take place.* *Mr. Martens added that it will be an exciting opportunity if we can have a demo - to identify from the webinar a couple of smaller clips that we can share via social media.* From dac6443842716558ed572d8af11e30f3ee88eb63 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 22:32:19 -0500 Subject: [PATCH 06/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 94 +++++++++++++++++++ 1 file changed, 94 insertions(+) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 0619bb2..fc83b28 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -5,6 +5,100 @@ layout: default title: Meeting Notes parent: Technical Committee --- +# Notes of the Technical Committee Meeting - April 14, 2023 +**The Meeting Began at 11:02am EDT.** + +**Attendees: J. Stine, N. Southern, T. Martens, D. Dahl, B. Epstein, C. Wüttke, E. Sewell, O. Coleman, S. Parlak, M. Cikan, A. Krish, U. Dincel, D. Larson.** + +**Notice of Recording - N. Southern** + +**Linux Foundation Anti-Trust Statement - N. Southern** + +## Basic Welcome to Meeting Attendees -- T. Martens ## + +## Review of Agenda, Opening Comments - C. Wüttke ## + +*Mr. Wüttke greeted and welcomed everyone and covered the following agenda items* + +* Director Updates - Q2 Deliverables & Member Engagement - J. Stine +* Technical Roadmap - Status & Updates - D. Dahl +* Trust Mark Initiative - Next Steps - O. Coleman +* Partnerships - Updates & Testbed Investments- J. Stine & T. Martens +* Comments, Questions and Miscellaneous - Group +* Closing Remarks - C. Wüttke + +## Director Updates - Q2 Deliverables & Member Engagement - J. Stine ## + +*Mr. Stine noted that today is the designated deadline for Q2 deliverables which need to be placed in issues/milestones in GitHub. He will be reviewing OVON's status on Q1 and emphasized the significant progress made by the Open Voice Team from January through March.* + +*Mr. Stine also noted that some things were deliberately not done this quarter, and these will be reviewed as well.* + +*Mr. Stine noted the great effort that is going into the question of member engagement and advancing broader participation on the project. Fulsome conversations have taken place - with Ms. Coleman, Mr. Snider, Mr. Pappas, and others - regarding how OVON onboards new individuals and connects with new individuals through the nonprofit's website. Efforts are being made to expand the Open Voice mailing list from 1200 to 1600-2000 contacts and reach more deeply into the conversational AI and voice ecosystems. This work is currently underway and numerous related face-to-face meetings will take place at Project Voice in Chattanooga, TN in ten days.* + +*Mr. Stine, Dr. Larson and Dr. Dahl are also proactively discussing ways to bring greater participation into the interoperability work group of the OVON Architecture Work Group. There have been 47 individuals invited to AWG meetings, but attendance averages between 8-10 persons per meeting, usually the same persons at every meeting. Slack has 100+ names. This raises critical questions about how to further engage, to bring more people in, and establish processes for participation that accept the fact that people may have conflicts during the week.* + +## Technical Roadmap - Status & Updates - D. Dahl ## + +*The Architecture Work Group has made significant progress in the past month, including some of the following highlights:* + +* Continue Identification and Authentication of Agents activity – requirements document finished/published at the end of April +* Additional work on Dialog Events – communicating natural language messages between users and systems. How is this represented in an interoperable way? + * Preliminary work in GitHub + * Gelling toward a solid specification +* Initial draft of spec for Delegation messages +* Interoperability Webinar including interoperability demo March 22 - at least 1-2 participants recruited in architecture work as product of this webinar +* Discussions with government of Estonia + +###### _A question was posed here in the chat, about whether or not we can authenticalte the content synthetically generated by an AI. Debbie pointed out that this is not one of the requirements. We cannot tell in many cases that the content was synthetically generated - so this will be a formidable problem if we expect a system to do this when people cannot._ ###### + +*Dr. Dahl also presented the current OVON roadmap that has been updated with a few items. The highlights are as follows:* + +* Demos and specifications are the primary work of the Interoperability Team. Delegation and dialog event specs are ongoing. +* Demos - AWG completed Demo #1 - a simple communication link +* Demo 2 - to be finished 5/18 - will include agent-agent passage of information, not just simple delegation. This is on track for completion by 5/18. +* The group has in the to-do column updated Demo #3, which will start to show how agents can discover each other and transmit information about the history of a conversation from one agent to another, so that the second agent has some idea of what was going on in the earlier discussion, plus contextual information. +* The architecture group has two new specs forthcoming: + * A 'Can You Handle It?' spec - where an assistant asks another if it can handle a request + * Another specification to help implementers of our work - an implementation hints document + +##### _Dr. Larson added that some of specifications will not result in publications but will be placed in the GitHub repository._ ##### + +## Welcome to New Attendees - J. Stine ## + +*Mr. Stine took time out, here, to introduce several new attendees to the Technical Committee meeting, and vice versa. Newcomers included Mr. Parlak, Mr. Cikan, Mr. Krish and Mr. Dincel.* + +*At this point, Dr. Dahl's discussion of the technical roadmap resumed.* + +##### _Dr. Larson asked Dr. Dahl for a systematic explanation and breakdown of what the colors mean, beyond the simple fact that they represent the various ongoing work streams; Dr. Dahl agreed, as an action item, to illustrate everything by creating a corresponding 'key,' that includes all of the representative icons._ ##### + +*Dr. Dahl then reviewed the Q2 Deliverables from the roadmap, delineated as follows:* +* Draft Interoperability Specifications/Delegation Patterns +* Second Demo +* Authentication of Agents Task Force +* Third Demo + +##### _Dr. Dahl mentioned that she has been over the milestones in Github and some work still needs to be done - certain milestones need to be split as they have been refined over time._ ##### + +##### _Mr. Sewell then asked the question will the test bed spec include infrastructure, authentication, data security and privacy considerations? Dr. Dahl deferred this question to the section of the discussion led by Mr. Martens and Mr. Stine._##### + +*Dr. Dahl concluded by discussing the continuance of execution of H1 goals - rewriting the specifications, and the demonstrations, and the prioritizations of these. She also stated that the Architecture Work Group is currently recruiting external partners; this will be discussed/reviewed later in the meeting - as will the discussion of specific use cases.* + +## Trust Mark Initiative - Oita Coleman ## + +*Ms. Coleman thanked everyone for their support on the launch a couple of weeks earlier - during the webinar, as well as various endorsements and messages of support. She noted that the launch was only the beginning and there is still much left to go.* + +*Ms. Coleman then presented the Trello Board - an illustration of everything that was accomplished to reach the Trust Mark launch date. These included the following:* +* Completing the code of ethics +* Sample module training course - 'voice data analytics' +* Training course e-learning design document +* Completion of Ethical Guidelines v. 2.0 white paper by March 24th +* Addition of Advisory Board Members +* Trust Mark Webinar on March 29th +* Completion of Self-Assessment Checklist Template +* Lance PR Rollout Plan + + + # Notes of the Technical Committee Meeting - March 10, 2023 **The Meeting Began at 10:07am EST.** From 5608da3837bcbe9b1d57dc19f7e20acdc03f9739 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 22:34:21 -0500 Subject: [PATCH 07/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index fc83b28..db2c626 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -43,11 +43,11 @@ parent: Technical Committee * Continue Identification and Authentication of Agents activity – requirements document finished/published at the end of April * Additional work on Dialog Events – communicating natural language messages between users and systems. How is this represented in an interoperable way? - * Preliminary work in GitHub - * Gelling toward a solid specification -* Initial draft of spec for Delegation messages -* Interoperability Webinar including interoperability demo March 22 - at least 1-2 participants recruited in architecture work as product of this webinar -* Discussions with government of Estonia + * Preliminary work in GitHub + * Gelling toward a solid specification + * Initial draft of spec for Delegation messages + * Interoperability Webinar including interoperability demo March 22 - at least 1-2 participants recruited in architecture work as product of this webinar + * Discussions with government of Estonia ###### _A question was posed here in the chat, about whether or not we can authenticalte the content synthetically generated by an AI. Debbie pointed out that this is not one of the requirements. We cannot tell in many cases that the content was synthetically generated - so this will be a formidable problem if we expect a system to do this when people cannot._ ###### @@ -58,8 +58,8 @@ parent: Technical Committee * Demo 2 - to be finished 5/18 - will include agent-agent passage of information, not just simple delegation. This is on track for completion by 5/18. * The group has in the to-do column updated Demo #3, which will start to show how agents can discover each other and transmit information about the history of a conversation from one agent to another, so that the second agent has some idea of what was going on in the earlier discussion, plus contextual information. * The architecture group has two new specs forthcoming: - * A 'Can You Handle It?' spec - where an assistant asks another if it can handle a request - * Another specification to help implementers of our work - an implementation hints document + * A 'Can You Handle It?' spec - where an assistant asks another if it can handle a request + * Another specification to help implementers of our work - an implementation hints document ##### _Dr. Larson added that some of specifications will not result in publications but will be placed in the GitHub repository._ ##### From 1721fb1a522fd7f476baf3dc686fc49f065f14d2 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:26:13 -0500 Subject: [PATCH 08/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 33 +++++++++++++++++-- 1 file changed, 31 insertions(+), 2 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index db2c626..63dbb6d 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -90,13 +90,42 @@ parent: Technical Committee *Ms. Coleman then presented the Trello Board - an illustration of everything that was accomplished to reach the Trust Mark launch date. These included the following:* * Completing the code of ethics * Sample module training course - 'voice data analytics' -* Training course e-learning design document +* Training course e-learning design document done with Linux Foundation * Completion of Ethical Guidelines v. 2.0 white paper by March 24th * Addition of Advisory Board Members * Trust Mark Webinar on March 29th -* Completion of Self-Assessment Checklist Template +* Completion of Self-Assessment Checklist/Maturity Model Template * Lance PR Rollout Plan +*Ms. Coleman noted that for Q2 2023, her group wants to have the following:* +* At least the bare bones of the training course available or completed by the end of June, +* The Self-Assessment and Maturity Model Phase 1 implementation by the end of June +* Participating organizations helping them complete the training, review the training, and review the self-assessment. +* Any design work on badges should also be done at that time. + +*Ms. Coleman then ran through a basic overview of the Trust Mark initiative for the benefit of newcomers to the meeting, and Mr. Stine stressed that the team plans to ask friend organizations of the Open Voice Network, including Sestek, to consider endorsing Trust Mark as it gets further along in development.* + +##### _Mr. Krish responded by expressing Sestek's enthusiasm for the initiative and strong desire to be involved. _##### + +##### _Ms. Coleman noted that she has heard from Mr. Wüttke, and his organization, Schwarz, has indicated its support of the Trust Mark program as well. Mr. Wüttke affirmed this in meeting and noted that it feels particularly critical and relevant to his team in light of the increasing commonality of Generative AI and ChatGPT._ ##### + +##### _Mr. Stine then informed Mr. Dincel that he will shortly forward pertinent information about Trust Mark to the Sestek team._ ##### + +##### _In the chat, Dr. Larson asked when E-learning will be available. Ms. Coleman stated that her team is working with the instructional design team of the Linux Foundation, and has submitted to them an instructional design document from which they are constructing a course. Dr. Larson stated that he is most interested in the content of the course. Ms. Coleman stated that the content - which is in active development - will be based on the Ethical Guidelines Document, Privacy Paper and the Security White Paper._ ##### + +*Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* + +By 4/14/2023 +* Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. +By 5/10/2023 +* 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. +* Work along with the developer to figure out implementation and internationalization requirements. +By 6/14/2023 +* Implementation phase working closely with the developer to implement the assessment as an interactive and multilingual tool. We would like to consider implementing it in multiple languages from the start, so we'll need to consider internationalization at this stage. +By 6/28/2023 +* 2 weeks of early adopter / intense testing to ensure functional requirements are met. +By 6/30/2023 +* Launch (based on feedback / testing). # Notes of the Technical Committee Meeting - March 10, 2023 From d61f88114f864923e92c6a14922968f860b2ed18 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:28:04 -0500 Subject: [PATCH 09/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 63dbb6d..9a67058 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -105,7 +105,7 @@ parent: Technical Committee *Ms. Coleman then ran through a basic overview of the Trust Mark initiative for the benefit of newcomers to the meeting, and Mr. Stine stressed that the team plans to ask friend organizations of the Open Voice Network, including Sestek, to consider endorsing Trust Mark as it gets further along in development.* -##### _Mr. Krish responded by expressing Sestek's enthusiasm for the initiative and strong desire to be involved. _##### +##### _Mr. Krish responded by expressing Sestek's enthusiasm for the initiative and strong desire to be involved._ ##### ##### _Ms. Coleman noted that she has heard from Mr. Wüttke, and his organization, Schwarz, has indicated its support of the Trust Mark program as well. Mr. Wüttke affirmed this in meeting and noted that it feels particularly critical and relevant to his team in light of the increasing commonality of Generative AI and ChatGPT._ ##### @@ -117,7 +117,7 @@ parent: Technical Committee By 4/14/2023 * Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. -By 5/10/2023 +By 5/10/2023 * 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. * Work along with the developer to figure out implementation and internationalization requirements. By 6/14/2023 @@ -128,6 +128,7 @@ By 6/30/2023 * Launch (based on feedback / testing). + # Notes of the Technical Committee Meeting - March 10, 2023 **The Meeting Began at 10:07am EST.** From de7a7fcce73dc1f80688d763889b152b90957503 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:29:45 -0500 Subject: [PATCH 10/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 9a67058..4f9f729 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -117,7 +117,7 @@ parent: Technical Committee By 4/14/2023 * Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. -By 5/10/2023 +By 5/10/2023 * 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. * Work along with the developer to figure out implementation and internationalization requirements. By 6/14/2023 @@ -126,6 +126,7 @@ By 6/28/2023 * 2 weeks of early adopter / intense testing to ensure functional requirements are met. By 6/30/2023 * Launch (based on feedback / testing). + From 901eec2588a744b412a34154641bdebe51935c9e Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:31:04 -0500 Subject: [PATCH 11/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 1 + 1 file changed, 1 insertion(+) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 4f9f729..d609e24 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -126,6 +126,7 @@ By 6/28/2023 * 2 weeks of early adopter / intense testing to ensure functional requirements are met. By 6/30/2023 * Launch (based on feedback / testing). + From 9d3ca30d7b3dd81174cca6bc7d7887d8c5c104a8 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:32:40 -0500 Subject: [PATCH 12/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 9 --------- 1 file changed, 9 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index d609e24..b17e709 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -119,15 +119,6 @@ By 4/14/2023 * Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. By 5/10/2023 * 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. -* Work along with the developer to figure out implementation and internationalization requirements. -By 6/14/2023 -* Implementation phase working closely with the developer to implement the assessment as an interactive and multilingual tool. We would like to consider implementing it in multiple languages from the start, so we'll need to consider internationalization at this stage. -By 6/28/2023 -* 2 weeks of early adopter / intense testing to ensure functional requirements are met. -By 6/30/2023 -* Launch (based on feedback / testing). - - From 2df8f463b51c0ba657359a733af89387aa6a1d7d Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:33:51 -0500 Subject: [PATCH 13/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index b17e709..7cba905 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -114,11 +114,9 @@ parent: Technical Committee ##### _In the chat, Dr. Larson asked when E-learning will be available. Ms. Coleman stated that her team is working with the instructional design team of the Linux Foundation, and has submitted to them an instructional design document from which they are constructing a course. Dr. Larson stated that he is most interested in the content of the course. Ms. Coleman stated that the content - which is in active development - will be based on the Ethical Guidelines Document, Privacy Paper and the Security White Paper._ ##### *Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* - By 4/14/2023 -* Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. -By 5/10/2023 -* 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. +By 5/12/2023 + From 1b88441ca00fff97dfca030d5e13e62ce4b13763 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:35:58 -0500 Subject: [PATCH 14/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 7cba905..9e83a80 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -114,10 +114,8 @@ parent: Technical Committee ##### _In the chat, Dr. Larson asked when E-learning will be available. Ms. Coleman stated that her team is working with the instructional design team of the Linux Foundation, and has submitted to them an instructional design document from which they are constructing a course. Dr. Larson stated that he is most interested in the content of the course. Ms. Coleman stated that the content - which is in active development - will be based on the Ethical Guidelines Document, Privacy Paper and the Security White Paper._ ##### *Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* -By 4/14/2023 -By 5/12/2023 - - +* By 4/14/23: Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. +* By 5/10/23 # Notes of the Technical Committee Meeting - March 10, 2023 From f0beacbdcc3aefe48ff6d4845bfa69aabc8d65fb Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:37:32 -0500 Subject: [PATCH 15/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 9e83a80..650bf4e 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -116,7 +116,8 @@ parent: Technical Committee *Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* * By 4/14/23: Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. * By 5/10/23 - + * 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. + * Work along with the developer to figure out implementation and internationalization requirements. # Notes of the Technical Committee Meeting - March 10, 2023 From a53d843b920b2f7e1008703567657194bebfb150 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 20 Apr 2023 23:40:00 -0500 Subject: [PATCH 16/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 650bf4e..b93287f 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -114,10 +114,16 @@ parent: Technical Committee ##### _In the chat, Dr. Larson asked when E-learning will be available. Ms. Coleman stated that her team is working with the instructional design team of the Linux Foundation, and has submitted to them an instructional design document from which they are constructing a course. Dr. Larson stated that he is most interested in the content of the course. Ms. Coleman stated that the content - which is in active development - will be based on the Ethical Guidelines Document, Privacy Paper and the Security White Paper._ ##### *Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* -* By 4/14/23: Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. +* By 4/14/23: + * Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. * By 5/10/23 * 4 weeks of reworking the questions and comparing with the 2 assessments we have been considering as models. * Work along with the developer to figure out implementation and internationalization requirements. +* By 6/14/2023: + * Implementation phase working closely with the developer to implement the assessment as an interactive and multilingual tool. We would like to consider implementing it in multiple languages from the start, so we'll need to consider internationalization at this stage. + * By 6/28/2023: 2 weeks of early adopter / intense testing to ensure functional requirements are met. + * By 6/30/2023 - Launch (based on feedback / testing). + # Notes of the Technical Committee Meeting - March 10, 2023 From c37f4611704a5f4f0c6466815df775ccc34018f7 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Fri, 21 Apr 2023 14:51:27 -0500 Subject: [PATCH 17/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 36 +++++++++++++++++-- 1 file changed, 34 insertions(+), 2 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index b93287f..5f7b8b8 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -103,7 +103,7 @@ parent: Technical Committee * Participating organizations helping them complete the training, review the training, and review the self-assessment. * Any design work on badges should also be done at that time. -*Ms. Coleman then ran through a basic overview of the Trust Mark initiative for the benefit of newcomers to the meeting, and Mr. Stine stressed that the team plans to ask friend organizations of the Open Voice Network, including Sestek, to consider endorsing Trust Mark as it gets further along in development.* +*Ms. Coleman then ran through a basic overview of the Trust Mark initiative for the benefit of newcomers to the meeting, and Mr. Stine stressed that the team plans to ask friend organizations of the Open Voice Network, including Sestek and Kaizen Voiz, to consider endorsing Trust Mark as it gets further along in development.* ##### _Mr. Krish responded by expressing Sestek's enthusiasm for the initiative and strong desire to be involved._ ##### @@ -113,7 +113,7 @@ parent: Technical Committee ##### _In the chat, Dr. Larson asked when E-learning will be available. Ms. Coleman stated that her team is working with the instructional design team of the Linux Foundation, and has submitted to them an instructional design document from which they are constructing a course. Dr. Larson stated that he is most interested in the content of the course. Ms. Coleman stated that the content - which is in active development - will be based on the Ethical Guidelines Document, Privacy Paper and the Security White Paper._ ##### -*Ms. Coleman then ran through the action item(s) and schedule associated with the Trust Mark Initiative, delineated as follows:* +*Ms. Coleman then ran through the action item(s) and schedule associated with the Self Assessment Maturity Model Web-Based Tool, delineated as follows. Each org that completes the assessment will have their own logon credentials and OVON will not have access to their results - this is for their own consideration and benefit* * By 4/14/23: * Work on the requirements for the implementation of the interactive assessment. This will result in a Statement of Work to recruit a developer resource. * By 5/10/23 @@ -123,6 +123,38 @@ parent: Technical Committee * Implementation phase working closely with the developer to implement the assessment as an interactive and multilingual tool. We would like to consider implementing it in multiple languages from the start, so we'll need to consider internationalization at this stage. * By 6/28/2023: 2 weeks of early adopter / intense testing to ensure functional requirements are met. * By 6/30/2023 - Launch (based on feedback / testing). +*Ms. Coleman then noted the basic structure of the cost breakdown - six weeks at fifteen hours a week, $100 an hour, which will equate to $9000.* +##### _Mr. Southern mentioned that his query to Christina Oliviero at the Linux Foundation soliciting help finding an external web developer is still active, and there should be additional news on this soon._ ##### +##### _Dr. Dahl inquired about the provision for maintaining this Trust Mark web app - it will be, she noted, a highly interactive application, and will need someone keeping an eye on it and updating it regularly._ ###### +##### _Ms. Coleman agreed with the need for this support. Some will come from Thrive Marketing. The need for additional support beyond this will be discussed, and will likely require funding. That will be scoped. There is a longer implementation plan of it having greater interactivity, and further guidance along the lines of 'if your score is a 7 out of 10, take these steps.' In other words, this is only Phase 1 implementation._ ##### +##### _Mr. Stine: any organization that asserts that it is an authority on ethical behavior in conversational AI lives in a glass house - so issues of privacy, data security and protection become doubly important for OVON. We need to doubly or triply check all of our own actions._ ##### +*Ms. Coleman noted that for this work to proceed, it will be taken to the Steering Committee at the end of this month for approval of allocated R&D dollars.* +##### _Mr. Stine: the Technical Committee has the right to approve this as coming from General Ledger number 6325._##### +*Ms. Coleman formally moved to have this allocation of $9000 approved by the technical community. Phase 1 development of the self-assessment maturity model tool. Mr. Stine gave it a second motion. There were no objections presented. Mr. Wüttke marked it duly approved, and Mr. Stine added one amendment - that any needed expenditures above and beyond the $9000 will come back to the Technical Committee.* +*Ms. Coleman: will work with Mr. Southern on identifying the correct/appropriate developer, creating the right implementation plan and making sure it falls within the above dollar amount. This proposal will be scoped for privacy/security and functionality and dollar allocation - the group must work within these constraints.* +*Said vote excludes the approval of additional and unrelated development monies for the website.* +## Investment and Co-Creation of the OVON Test Bed - J Stine and T. Martens ## +*Mr. Martens provided a quick overview of next steps for the Open Voice Network Test Bed and how the project plans to invest in components that can either enable or facilitate implementation of OVON specs together with Open Voice's partners. This is a watershed moment for Open Voice. On this note, the committee needs to alert this forum about upcoming requests for funding, and said committee will ultimately need to approve funding requests. Participants in the working groups will ultimately need to submit proposals that require funding and that can be included in the Test Bed discussion. Open Voice's relationship with the Estonian government and its work in voice AI will provide a viable use case; additional use cases will be sought in vertical sectors such as banking and insurance and retail and automotive. The goal will be not to run these kinds of implementations ourselves, but to invest strategically in partnerships and be an enabler of these use cases that will implement and exhibit the Open Voice specifications.* +*Mr. Martens stressed the need to clear this with the committee.* +##### _Dr. Dahl asked, should we say 'prototype' rather than 'MVP' which implies a product?_ ##### +##### _To this, Mr. Martens responded, 'MVP' actually implies a functionality we can showcase that has a benefit to the participating parties. However 'MVP' and 'prototype' do imply two sides of the same coin. This wording must be straightened out semantically. _##### +##### _Mr. Stine spoke directly to Sestek participants at this point, noting that in the months ahead, Open Voice Network will be actively recruiting firms such as Sestek, and rhetorically asked these participants if Sestek has enterprise customers - banking, insurance, retail, automotive, other verticals - who may be ready to begin expanding their voice, with their customers, outside the firewall, connecting to other partners that they have or to voice AI-welcoming and -using consumers? OVON would welcome this conversation and there would be reciprocal benefits for said organizations._ ##### +##### _Mr. Sewell asked Mr. Stine about his vision for the Open Voice Network Test Bed, and if he envisions it as more of a Proof of Concept or a lab test bed, or as more of a pilot Test Bed that could serve as a business opportunity, get customers in, and operationalize and scale?_ ##### +##### _Mr. Stine ultimately deferred this decision to the Technical Committee as a whole, but his instinct is that Open Voice would follow each of these steps in succession (1-2) and would want to get to the level of showing the business benefit in a production pilot as opposed to a lab POC. However, this is still to be determined and confirmed._ ##### +##### _Mr. Sewell then asked Mr. Stine if such a POC would be driving toward a standard._ ##### +##### _Mr. Stine said yes, without question. He foresees a testing process of the draft specifications coming out of the Architecture Working Group._ ##### +##### _Mr. Martens and others agreed, and further clarified/outlined that with the various OVON working groups, the specifications have been reviewed by multiple stakeholders; soon it will be time to test them under live conditions. He foresees two phases, starting with implementations that will overlap; if all parties agree that there is a benefit of rolling them out, they will be rolled out. Open Voice sees its role as discovering use cases and then facilitating them._ ##### +##### _Dr. Dahl expressed a concern about OVON getting into a position of offering a component (say a speech recognition component) as something OVON creates, branded with Open Voice, that could be directly implemented into a given product. Open Voice lacks the bandwidth for this kind of work. This would be out of scope for Open Voice and extremely expensive. It means being cautious about how Open Voice positions itself with individual components_ ##### +## Estonia ISA - Open Voice Network Collaborative Partnership ## +*Mr. Stine announced the steps that need to be taken with Estonian Information System Authority and OVON's colleagues therein. They wish to partner with The Open Voice Network in testing and publicly demonstrating OVON Interoperability Specifications. This fits within the long-range roadmap of the services that Bürokratt will provide for the citizens and different government functions of Estonia. Estonia sees itself also as publicly endorsing the Trust Mark, are exciting by its direction, and on this level they are eager to move forward. The following steps will need to take place:* +* Mr. Stine will meet shortly with Dr. Larson and Dr. Dahl to discuss a draft of the Interoperability Proposal for Estonia. +* Mr. Stine will be compiling a non-binding Letter of Intent for Estonia, using the standard Linux Foundation format as approved by Scott Nicholas and others and identify areas of collaboration, general deliverables, resource commitments and communication commitments. This will not be a binding contract. +* Mr. Stine also emphasized Open Voice's interest in exploring the same opportunities and partnerships with Sestek and Kaizen Voiz. +## Minutes Approval - Prior Technical Committee Meeting from Mar. 10th ## +*Mr. Stine put forth a first motion for approval, Ms. Coleman a second, and with no further objections, Mr. Southern marked the minutes duly approved. +**Adjournment - The Meeting formally adjourned at 12pm EDT.** + + # Notes of the Technical Committee Meeting - March 10, 2023 From 20cfb4f4ae9eb3ba5aedb76c718a739a566e6374 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Fri, 21 Apr 2023 14:53:40 -0500 Subject: [PATCH 18/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 18 +++++++++++++++--- 1 file changed, 15 insertions(+), 3 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 5f7b8b8..9fce1a9 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -137,21 +137,33 @@ parent: Technical Committee *Mr. Martens provided a quick overview of next steps for the Open Voice Network Test Bed and how the project plans to invest in components that can either enable or facilitate implementation of OVON specs together with Open Voice's partners. This is a watershed moment for Open Voice. On this note, the committee needs to alert this forum about upcoming requests for funding, and said committee will ultimately need to approve funding requests. Participants in the working groups will ultimately need to submit proposals that require funding and that can be included in the Test Bed discussion. Open Voice's relationship with the Estonian government and its work in voice AI will provide a viable use case; additional use cases will be sought in vertical sectors such as banking and insurance and retail and automotive. The goal will be not to run these kinds of implementations ourselves, but to invest strategically in partnerships and be an enabler of these use cases that will implement and exhibit the Open Voice specifications.* *Mr. Martens stressed the need to clear this with the committee.* ##### _Dr. Dahl asked, should we say 'prototype' rather than 'MVP' which implies a product?_ ##### -##### _To this, Mr. Martens responded, 'MVP' actually implies a functionality we can showcase that has a benefit to the participating parties. However 'MVP' and 'prototype' do imply two sides of the same coin. This wording must be straightened out semantically. _##### +##### _To this, Mr. Martens responded, 'MVP' actually implies a functionality we can showcase that has a benefit to the participating parties. However 'MVP' and 'prototype' do imply two sides of the same coin. This wording must be straightened out semantically._ ##### ##### _Mr. Stine spoke directly to Sestek participants at this point, noting that in the months ahead, Open Voice Network will be actively recruiting firms such as Sestek, and rhetorically asked these participants if Sestek has enterprise customers - banking, insurance, retail, automotive, other verticals - who may be ready to begin expanding their voice, with their customers, outside the firewall, connecting to other partners that they have or to voice AI-welcoming and -using consumers? OVON would welcome this conversation and there would be reciprocal benefits for said organizations._ ##### + ##### _Mr. Sewell asked Mr. Stine about his vision for the Open Voice Network Test Bed, and if he envisions it as more of a Proof of Concept or a lab test bed, or as more of a pilot Test Bed that could serve as a business opportunity, get customers in, and operationalize and scale?_ ##### + ##### _Mr. Stine ultimately deferred this decision to the Technical Committee as a whole, but his instinct is that Open Voice would follow each of these steps in succession (1-2) and would want to get to the level of showing the business benefit in a production pilot as opposed to a lab POC. However, this is still to be determined and confirmed._ ##### + ##### _Mr. Sewell then asked Mr. Stine if such a POC would be driving toward a standard._ ##### + ##### _Mr. Stine said yes, without question. He foresees a testing process of the draft specifications coming out of the Architecture Working Group._ ##### + ##### _Mr. Martens and others agreed, and further clarified/outlined that with the various OVON working groups, the specifications have been reviewed by multiple stakeholders; soon it will be time to test them under live conditions. He foresees two phases, starting with implementations that will overlap; if all parties agree that there is a benefit of rolling them out, they will be rolled out. Open Voice sees its role as discovering use cases and then facilitating them._ ##### + ##### _Dr. Dahl expressed a concern about OVON getting into a position of offering a component (say a speech recognition component) as something OVON creates, branded with Open Voice, that could be directly implemented into a given product. Open Voice lacks the bandwidth for this kind of work. This would be out of scope for Open Voice and extremely expensive. It means being cautious about how Open Voice positions itself with individual components_ ##### ## Estonia ISA - Open Voice Network Collaborative Partnership ## + *Mr. Stine announced the steps that need to be taken with Estonian Information System Authority and OVON's colleagues therein. They wish to partner with The Open Voice Network in testing and publicly demonstrating OVON Interoperability Specifications. This fits within the long-range roadmap of the services that Bürokratt will provide for the citizens and different government functions of Estonia. Estonia sees itself also as publicly endorsing the Trust Mark, are exciting by its direction, and on this level they are eager to move forward. The following steps will need to take place:* + * Mr. Stine will meet shortly with Dr. Larson and Dr. Dahl to discuss a draft of the Interoperability Proposal for Estonia. * Mr. Stine will be compiling a non-binding Letter of Intent for Estonia, using the standard Linux Foundation format as approved by Scott Nicholas and others and identify areas of collaboration, general deliverables, resource commitments and communication commitments. This will not be a binding contract. -* Mr. Stine also emphasized Open Voice's interest in exploring the same opportunities and partnerships with Sestek and Kaizen Voiz. + +*Mr. Stine also emphasized Open Voice's interest in exploring the same opportunities and partnerships with Sestek and Kaizen Voiz. + ## Minutes Approval - Prior Technical Committee Meeting from Mar. 10th ## -*Mr. Stine put forth a first motion for approval, Ms. Coleman a second, and with no further objections, Mr. Southern marked the minutes duly approved. + +*Mr. Stine put forth a first motion for approval, Ms. Coleman a second, and with no further objections, Mr. Southern marked the minutes duly approved.* + **Adjournment - The Meeting formally adjourned at 12pm EDT.** From 9a9c19277c5ea3c5f518899cccd4714a822b42d5 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Fri, 21 Apr 2023 14:54:32 -0500 Subject: [PATCH 19/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 9fce1a9..b88fbdc 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -158,7 +158,7 @@ parent: Technical Committee * Mr. Stine will meet shortly with Dr. Larson and Dr. Dahl to discuss a draft of the Interoperability Proposal for Estonia. * Mr. Stine will be compiling a non-binding Letter of Intent for Estonia, using the standard Linux Foundation format as approved by Scott Nicholas and others and identify areas of collaboration, general deliverables, resource commitments and communication commitments. This will not be a binding contract. -*Mr. Stine also emphasized Open Voice's interest in exploring the same opportunities and partnerships with Sestek and Kaizen Voiz. +*Mr. Stine also emphasized Open Voice's interest in exploring the same opportunities and partnerships with Sestek and Kaizen Voiz.* ## Minutes Approval - Prior Technical Committee Meeting from Mar. 10th ## From f9d97d57fa556140d73d58839bc16444043842c2 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Fri, 21 Apr 2023 15:16:55 -0500 Subject: [PATCH 20/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 425844c..6ffb208 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -25,7 +25,7 @@ parent: Technical Committee * Trust Mark Initiative - Next Steps - O. Coleman * Partnerships - Updates & Testbed Investments- J. Stine & T. Martens * Comments, Questions and Miscellaneous - Group -* Closing Remarks - C. Wüttke +* Closing Remarks - C. Wüttke ## Director Updates - Q2 Deliverables & Member Engagement - J. Stine ## From 4e0b43d0ebdfe23f57a592c51a7876756d368206 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Wed, 14 Jun 2023 14:55:38 -0500 Subject: [PATCH 21/35] Create eventsspecifications --- specs/eventsspecifications | 1 + 1 file changed, 1 insertion(+) create mode 100644 specs/eventsspecifications diff --git a/specs/eventsspecifications b/specs/eventsspecifications new file mode 100644 index 0000000..8b13789 --- /dev/null +++ b/specs/eventsspecifications @@ -0,0 +1 @@ + From f94cc502ffef7c1d9ff5cc22c9b0e60884b7b7fb Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Wed, 14 Jun 2023 20:18:42 -0500 Subject: [PATCH 22/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 93 +++++++++++++++++++ 1 file changed, 93 insertions(+) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 6ffb208..a7d9066 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -5,6 +5,99 @@ layout: default title: Meeting Notes parent: Technical Committee --- +ments for mechanisms to monitor privacy and ethical principles. +* Define requirements for mechanisms to evaluate trustworthiness of voice assistants. +* Review and improve Interoperability Patterns API (Delegation, Channeling and Mediation) and make it worthy for publication. + +*This gives people the opportunity to support Open Voice by performing specific tasks.* +*Dr. Larson and Dr. Dahl start the process with each individual by having a conversation with them about what their interests are, and if it's appropriate for them to start working with OVON* +*If it is appropriate, the team will try to find something of mutual benefit to that individual and to Open Voice.* +*Those who fill out the form - their information is sent to Evan via email.* + +##### _At this juncture, Mr. Stine pulled up the Open Voice Network website, which includes a drop-down list that gives various options for the many ways in which one can contribute to/participate in The Open Voice Network. This is still as of yet a work in progress. It lists tasks and skills - if you have an interest in x, y, z, etc. This is a new observation to the Open Voice Network website._ ##### + + +## Recruitment of Demonstrators - J. Stine ## + +*Mr. Stine has had informal conversations with six organizations, which he then listed in the meeting and asked to be kept confidential - plus one major voice assistant based in Europe. At least five of the prospects have asked to know more; three of those five have bene unusually enthusiastic, but collaboration will require resources, time and a product that is not in a lab but operable and has been implemented.* + +*Mr. Stine and Dr. Dahl will be working together on these prospective partnerships. Heath care may be a particularly strong focus, as it mirrors one of Rainer Türner (Estonia ISA)'s use cases - specifically, making an appointment with a pediatrician. The Estonian ISA plans to work with three specific use cases - two Estonian and one English. They expect at least one use case from Open Voice that still needs to be pinned down.* + +##### _Dr. Larson pointed out that recruitment of demonstrators may involve Estonia but not necessarily, and they must be coordinated carefully and discussed with Dr. Dahl prior to onboarding them._##### + +##### _Mr. Stine pointed out that this will require resources of the collaborative engagement party._ ##### + +## Trust Mark Initiative Update - O. Coleman ## + +*The Trustmark team is actively working on delivery the self-assessment tool and training course.* + +*Regarding the self-assessment tool: the developer (Luis) has been interviewed and hired under contract as a vendor to Linux Foundation, and began his work on Monday. yesterday there was a meeting to go over the requirements and he showed Oita, Valeria and Nathan his design mock-ups; that meeting went very well; he will now begin development; he's planning to deliver a version of the tool the week of June 19th, when the initial testing will begin.* + +*Overall the goal for wrapping up the project is for initial release to early adopters by the end of July.* + +*Luis is very easy to work with and extremely intelligent and capable, and his work so far has been extremely forward-thinking. The tool - the questions - are text based and text intensive; on his own, Luis has added an admin tool so that the Open Voice team can do edits themselves as needed and won't need to go to a web developer. It doesn't involve updates to the code, but to metadata behind the questions only.* + +*The Trustmark training course is being actively worked out. The content draft will be completed by the end of the day today, and content revisions and copy editing are slated to begin by Monday June 12, 2023. Lotas Productions (Jim Kenneally) has been critical in this regard. They will be providing voice talent - videos and voiceovers, including synthetic voices, for some of the images.* + +*In approximately one week, the content will be handed over to the Linux Foundation to be uploaded into the LMS System. The launch date is projected to be tentatively August 2, 2023. Linux Foundation will provide a more definitive date once they see the content.* + +*A link to the draft chapter information is posted on Slack (ethics channel of OVON) in google doc form for anyone who wants to review in this format (pre-copy editing) and provide feedback.* + +###### _Mr. Stine has noted by way of elaboration that Lotas Productions, run by Jim Kenneally, has evinced strong support of the Trust Mark Initiative; his firm procures and vets voice talent for leading NYC ad agencies. He's contributing pro bono to bring in top-tiered voice talent to be voicing the training courses for the Open Voice Network._ ##### + +##### _ Dr. Dahl asked if there would be a benefit of briefly acknowledging Trust Mark in the Interoperability Webinar. Ms. Coleman and others strongly agreed. This will take the form of an announcement._ ##### + +##### _Mr. Stine noted that with respect to the Trustmark Initiative, Open Voice Network is looking for organizations to endorse its principles. It is also open to individuals. If you are able to act as an individual without committing your company, that's a place to show your name. Accordingly, several names are on the website in this capacity. Every company is different, and whether or not one can endorse as an individual, it isn't true in every case. Open Voice Network is looking to the leaders of the Open Voice Network to endorse the initiative as individuals._ ##### + +## New Terms for the OVON Glossary - Dr. Larson ## + +*Many discussions have taken place between Dr. Larson and numerous people involving controversial terms that follow:* + +* Conversational Assistant +* Interoperability Pattern +* Native Pattern +* Delegation Interoperability Pattern +* Mediation Interoperability Pattern +* Channeling Interoperability Pattern +* Dialog Event +* Dialog Event Feature +* Assistant Identity +* Assistant Authentication +* Tampered Assistant +* Assistant Browser +* Endpoint + +*Accordingly Dr. Larson put forward revised definitions of each of the above in this meeting. He proposed that these terms be placed in the Open Voice Network Glossary and that the Technical Committee vote on them, and that any comments be reviewed in the next Technical Committee Meeting.* + +##### _Mr. Stine raised a key concern about Dr. Larson's proposed definition of the term 'Assistant Browser' - on the basis of the presumption that this is synonymous with 'Assistant.'_ ##### + +##### _Dr. Larson disagreed and noted that an Assistant Browser is a special software module that manages the speaker, the display, the keyboard, that accepts requests from user(s) and sends them back. This, Dr. Larson noted, is different from an 'assistant.'_ ##### + +*Dr. Larson put this approval forward as a first motion. Mr. Stine put forward a second motion. There were no objections. Said definitions were designated duly approved.* + +## Closing Remarks - Group ## + +##### _Moving forward, Dr. Larson asked that the Technical Committee slides be sent out a couple of days before the meeting, so that he has more time to respond._ ##### +##### _Ms. Coleman: the Privacy & Security Work Group has renewed its efforts to review what is now happening in the privacy/security sphere from a regulatory/legislative perspective, globally. It is evaluating how to update its policy white papers accordingly. This week it will be discussing the upcoming EU AI Act. It has passed committee, and will be up for a formal vote between June 12-15. They have defined risk categories, from low to unacceptable. Any AI-based product services that do biometric processing are deemed unacceptable risks. Some friends of OVON may be doing work in this category; it raises related questions for Open Voice and how OVON should respond. The bottom line is that Open Voice needs to understand what the implications are. It may be worth having a 10 minute discussion in the next Technical Committee Meeting on the outcome of this vote and what steps should be taken accordingly._ ##### + +## Adjournment - With no additional comments or areas of discussion, this meeting was adjourned at 11:54am Eastern. ## + + +## Action Items - Group ## + +* J. Stine, D. Dahl, T. Martens - determine who can be brought into the Estonia project from Open Voice's internal network, as contributors.* +* D. Dahl: add a list of authors and a document summary and AWG (group) acknowledgments to the Dialog Event Specification Paper. +* D. Dahl: get final slides for interoperability webinar to N. Southern by Monday +* J. Stine & group: after the 6/15 webinar send out solicitations to participants to volunteer. +* N. Southern: continue to work on sourcing emails not yet available in Hubspot. +* Group members: if you know of any friends who might be interested in volunteering for Open Voice Network, point them to the drop-down on the OVON website. +* D. Dahl & Estonia team - pinpoint one specific use case to do with Estonia. +* J. Stine - track down and review job description sent to him by Dr. Larson and Dr. Dahl +* O. Coleman -provide Trust Mark sentence for the Interoperability Webinar. + + + + # Notes of the Technical Committee Meeting - April 14, 2023 **The Meeting Began at 11:02am EDT.** From 6333ab29cee24d9ae890f4a6765e15f899457fc0 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Wed, 14 Jun 2023 20:19:29 -0500 Subject: [PATCH 23/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 102 +++++++++++++++++- 1 file changed, 101 insertions(+), 1 deletion(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index a7d9066..6b96048 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -5,7 +5,91 @@ layout: default title: Meeting Notes parent: Technical Committee --- -ments for mechanisms to monitor privacy and ethical principles. +# Notes of the Open Voice Network Technical Committee Meeting - June 9, 2023 +**The Meeting Began at 11:01am EDT.** + +**Attendees: J. Stine, N. Southern, O. Coleman, T. Martens, J. Larson, D. Dahl, B. Epstein, H. Pappas** + +**Basic Welcome to Meeting Attendees -- T. Martens** + +**Notice of Recording - T. Martens** + +**Reading of Linux Foundation Anti-Trust Statement - N. Southern** + +## Minutes Approval of May 12, 2013 Technical Committee Meeting - T. Martens ## + +*Mr. Stine put forth a first approval motion; Dr. Larson seconded that motion; Mr. Southern marked the minutes duly approved.* + +## Review of Agenda, Opening Comments - T. Martens ## + +*Mr. Martens noted the following agenda items for today's meeting* + +* Bürokratt update on Estonia - Working Plan - D. Dahl +* Interoperability Roadmap update - D. Dahl +* Interoperability Webinar on June 15th - D. Dahl +* Recruitment for Demonstrators- J. Stine & T. Martens +* Trustmark Initiative Update/Self-Assessment Tool - O. Coleman +* New Terms for OVON Glossary - J. Larson +* Miscellaneous Comments/Questions - Group +* Closing Remarks - T. Martens + +## Bürokratt update on Estonia - Working Plan - D. Dahl ## + +*Dr. Dahl just finished a call with Rainer Türner from the Estonia ISA and that great progress is made overall.* + +*Mr. Türner is creating a series of user stories. For the first one, Bürokratt wants to provide voice-based communication based on OVON specifications to integrate their technology with third-party participants following the same specifications. This is the overarching goal.* + +*The Estonia ISA plans to ultimately implement OVON specifications in their Bürokratt system. and provide feedback to the Open Voice Network team on what they find useful or not useful; they are willing to do this and assist OVON with comments on aspects of the initiative for which the details still need to be worked out.* + + + +*Reciprocally, OVON needs to provide specifications and documentation and help them with any issues that they have surrounding the specifications.* + +*Rainer Türner and Dr. Debbie Dahl are the two technical project leads.* + +*A full project plan can be found at the url https://github.com/orgs/buerokratt/projects/35. The goal is to have everything wrapped up by December 31, 2023.* + +##### _Mr. Stine asked what Dr. Dahl envisions to be the incrementally additional resources for the Open Voice Network to make this happen? What else is in front of us, besides recruitment and involving third-party practitioners?_ ##### + +##### _Dr. Dahl noted that interoperability specifications are not nearly as far along as the dialog events specifications, and Mr. Türner is very willing to help with solidifying these specs, but any other OVON people willing to focus on this would be helpful. Emmett Coin is already deeply involved because he's implementing it. He's searching for gaps and things to tighten up/resolve. The goal: Bürokratt is doing the heavy lifting and providing expertise and advice._ ##### + +##### _Dr. Dahl then brought said project plan up on the screen._ ##### + +##### _Mr. Martens noted that Estonia has implemented the current version of Bürokratt using Raza, and if they start using OVON specifications, it would be a welcome opportunity for other OVON members/contributors to come in and provide their technologies using OVON specifications. This call is welcome/open to everyone. OVON will act as facilitators._ ##### + +*Dr. Dahl: a couple of other things are going on in parallel focusing on the interoperability effort with Estonia. One is the overarching collaboration agreement; the agreement that Mr. Stine put together also includes a privacy and security component.* + +##### _Mr. Stine noted that an overarching piece of handshake documentation has been extended to Raïner's supervisor, Kaupo Laagriküll. This agreement speaks to two engagements: interoperability, and the trial and testing of trustmark educational and self-assessment tools as they become ready. Part of it will be led not only by Laagriküll but by the overarching leader of the Estonia ISA., Ott Velsberg. There is strong interest in testing, using, exploring the Trustmark Tools Ms. Coleman and co. are developing, and we are to come back to Estonia as those tools are ready for testing and review. The agreement is in Kaupo's hands and he will circle back to Mr. Stine soon with thoughts._ ##### + +##### _It was clarified that the second half of the Estonia partnership concerns Trustmark, in lieu of Privacy and Security._ ##### + +*In terms of the Estonian partnership: a draft outreach email that would be going to potential third party participants with the Estonian initiative. There are six targeted companies (prospects) and Mr. Stine will be reviewing these prospects with Dr. Dahl before the emails go out.* + +## Interoperability Roadmap Update - D. Dahl ## + +* The Dialog Events Specification v.1 draft - has been edited/formatted; ready to be published although it needs to include a list of the authors' names. +* The Interoperability Webinar is slated for Thu. Jun. 15th - it will have a panel format, and will include the announcement of a dialog event specification. +* Dr. Dahl and Dr. Larson have done considerable work on identifying volunteer opportunities, and there are six different items. It has been useful wrt pointing new contributors in the right direction. +* Dr. Larson is also maintaining a log of people who have expressed some level of interest in contributing to Open Voice. + +##### _It was also suggested that all the members of the Architecture Work Group be asked to sign off on the Dialog Events Specification document, but noted in response that the AWG includes over 100 people, so it would probably make more sense to save this for the next publication._ ##### + +## Interoperability Webinar on June 15th - 12pm Eastern - D. Dahl - Highlights ## +* Announcement of Publication of Dialog Events Specification 1.0 +* Demonstration of simple negotiation between assistants and passing on dialog history +* Round Table Discussion of Demo Features + +##### _Mr. Stine noted thirty registrants as of Thursday, June 8th, asked anyone to register who has not done so, and noted two more rounds of promotion yet to go out. The number of existing registrations should escalate dramatically in the last 24 hours, so a total registration of 60-70 persons is projected. Mr. Stine noted a steady increase in webinar attendance numbers over the past months, because Ovon has increased the size of its mailing list from roughly 1200 individuals to around 3300 individuals. ##### + +## Webinar - Recruitment for Demonstrators - J. Stine and T. Martens ## + +*Dr. Stine and Dr. Dahl have opened up a new series of volunteer opportunities for Open Voice Network contributors, in the following key areas:* + +* Write OVON-based assistant with minimal functionality that can be used to test other assistants using OVON protocols. +* Refine and develop a Python library for Dialog Event processing. +* Create usage scenarios to measure progress against the development of functionality. +* Extend and refine dialog history and context representations. +* Define requirements for mechanisms to monitor privacy and ethical principles. * Define requirements for mechanisms to evaluate trustworthiness of voice assistants. * Review and improve Interoperability Patterns API (Delegation, Channeling and Mediation) and make it worthy for publication. @@ -98,6 +182,22 @@ ments for mechanisms to monitor privacy and ethical principles. + +## Action Items - Group ## + +* J. Stine, D. Dahl, T. Martens - determine who can be brought into the Estonia project from Open Voice's internal network, as contributors.* +* D. Dahl: add a list of authors and a document summary and AWG (group) acknowledgments to the Dialog Event Specification Paper. +* D. Dahl: get final slides for interoperability webinar to N. Southern by Monday +* J. Stine & group: after the 6/15 webinar send out solicitations to participants to volunteer. +* N. Southern: continue to work on sourcing emails not yet available in Hubspot. +* Group members: if you know of any friends who might be interested in volunteering for Open Voice Network, point them to the drop-down on the OVON website. +* D. Dahl & Estonia team - pinpoint one specific use case to do with Estonia. +* J. Stine - track down and review job description sent to him by Dr. Larson and Dr. Dahl +* O. Coleman -provide Trust Mark sentence for the Interoperability Webinar. + + + + # Notes of the Technical Committee Meeting - April 14, 2023 **The Meeting Began at 11:02am EDT.** From 3205c1aa08f027de69f3c6008b0091c3b963de13 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Wed, 14 Jun 2023 20:20:47 -0500 Subject: [PATCH 24/35] Update technical-committee-meeting-notes.md --- .../technical-committee-meeting-notes.md | 17 +---------------- 1 file changed, 1 insertion(+), 16 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 6b96048..92d3302 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -35,7 +35,7 @@ parent: Technical Committee ## Bürokratt update on Estonia - Working Plan - D. Dahl ## -*Dr. Dahl just finished a call with Rainer Türner from the Estonia ISA and that great progress is made overall.* +*Dr. Dahl just finished a call with Rainer Türner from the Estonia ISA and noted that great progress is being made overall.* *Mr. Türner is creating a series of user stories. For the first one, Bürokratt wants to provide voice-based communication based on OVON specifications to integrate their technology with third-party participants following the same specifications. This is the overarching goal.* @@ -183,21 +183,6 @@ parent: Technical Committee -## Action Items - Group ## - -* J. Stine, D. Dahl, T. Martens - determine who can be brought into the Estonia project from Open Voice's internal network, as contributors.* -* D. Dahl: add a list of authors and a document summary and AWG (group) acknowledgments to the Dialog Event Specification Paper. -* D. Dahl: get final slides for interoperability webinar to N. Southern by Monday -* J. Stine & group: after the 6/15 webinar send out solicitations to participants to volunteer. -* N. Southern: continue to work on sourcing emails not yet available in Hubspot. -* Group members: if you know of any friends who might be interested in volunteering for Open Voice Network, point them to the drop-down on the OVON website. -* D. Dahl & Estonia team - pinpoint one specific use case to do with Estonia. -* J. Stine - track down and review job description sent to him by Dr. Larson and Dr. Dahl -* O. Coleman -provide Trust Mark sentence for the Interoperability Webinar. - - - - # Notes of the Technical Committee Meeting - April 14, 2023 **The Meeting Began at 11:02am EDT.** From f78aa7ab9f1c2d040da77ead8a23bfe65f012545 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Wed, 14 Jun 2023 20:22:19 -0500 Subject: [PATCH 25/35] Update technical-committee-meeting-notes.md --- technical-committee/technical-committee-meeting-notes.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/technical-committee/technical-committee-meeting-notes.md b/technical-committee/technical-committee-meeting-notes.md index 92d3302..d517e3c 100644 --- a/technical-committee/technical-committee-meeting-notes.md +++ b/technical-committee/technical-committee-meeting-notes.md @@ -107,7 +107,7 @@ parent: Technical Committee *Mr. Stine and Dr. Dahl will be working together on these prospective partnerships. Heath care may be a particularly strong focus, as it mirrors one of Rainer Türner (Estonia ISA)'s use cases - specifically, making an appointment with a pediatrician. The Estonian ISA plans to work with three specific use cases - two Estonian and one English. They expect at least one use case from Open Voice that still needs to be pinned down.* -##### _Dr. Larson pointed out that recruitment of demonstrators may involve Estonia but not necessarily, and they must be coordinated carefully and discussed with Dr. Dahl prior to onboarding them._##### +##### _Dr. Larson pointed out that recruitment of demonstrators may involve Estonia but not necessarily, and they must be coordinated carefully and discussed with Dr. Dahl prior to onboarding them._ ##### ##### _Mr. Stine pointed out that this will require resources of the collaborative engagement party._ ##### @@ -129,7 +129,7 @@ parent: Technical Committee ###### _Mr. Stine has noted by way of elaboration that Lotas Productions, run by Jim Kenneally, has evinced strong support of the Trust Mark Initiative; his firm procures and vets voice talent for leading NYC ad agencies. He's contributing pro bono to bring in top-tiered voice talent to be voicing the training courses for the Open Voice Network._ ##### -##### _ Dr. Dahl asked if there would be a benefit of briefly acknowledging Trust Mark in the Interoperability Webinar. Ms. Coleman and others strongly agreed. This will take the form of an announcement._ ##### +##### _Dr. Dahl asked if there would be a benefit of briefly acknowledging Trust Mark in the Interoperability Webinar. Ms. Coleman and others strongly agreed. This will take the form of an announcement._ ##### ##### _Mr. Stine noted that with respect to the Trustmark Initiative, Open Voice Network is looking for organizations to endorse its principles. It is also open to individuals. If you are able to act as an individual without committing your company, that's a place to show your name. Accordingly, several names are on the website in this capacity. Every company is different, and whether or not one can endorse as an individual, it isn't true in every case. Open Voice Network is looking to the leaders of the Open Voice Network to endorse the initiative as individuals._ ##### From bec5b4f2306ea6f05490d3a6fe5cc59287bba1c1 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 14:04:40 -0500 Subject: [PATCH 26/35] Create Interoperability of Conversational Assistants --- ...eroperability of Conversational Assistants | 698 ++++++++++++++++++ 1 file changed, 698 insertions(+) create mode 100644 specs/Interoperability of Conversational Assistants diff --git a/specs/Interoperability of Conversational Assistants b/specs/Interoperability of Conversational Assistants new file mode 100644 index 0000000..516cd7d --- /dev/null +++ b/specs/Interoperability of Conversational Assistants @@ -0,0 +1,698 @@ + +2022.08.12 +Version 1.0-B +Interoperability of Conversational Assistants A NEW APPROACH +The Open Voice Network +Architecture Work Group of the Technical Committee +12 August 2022 + + +Executive Summary +The Open Voice Network (OVON) is an open-source community of the Linux Foundation, dedicated to developing the technical standards and usage guidelines for the emerging world of conversational Artificial Intelligence and, within that, voice assistance. +The future of conversational Artificial Intelligence is a future of diversity – of infrastructure providers and platforms, of enabling technologies, devices, and enterprise use cases. Conversational AI – and within conversational AI, voice assistance -- will be a primary interface to the internet, personal transportation, smart environments (domestic and enterprise), and the immersive digital world of gaming and augmented and virtual experiences. +To reap the greatest economic and social value of conversational assistance, we believe that voice must work like the web, enabling users to access any voice-enabled content destination regardless of platform. Interoperability – the sharing of dialogs between conversational assistants and agents of different technological parentage – is an essential capability to reach the value of what will become a Worldwide Voice Web. It must also be worthy of individual and enterprise user trust; readers will note references here to OVON work in data privacy, data security, and ethical use. +This paper is a step on the path to conversational assistance interoperability. In this paper, we assert that the proper model for conversational assistance interoperability is a human one – and suggest that humans generally resolve questions with one another through a process of either mediation or delegation. Using this framework, we identify the jobs to be done for priority constituencies, the architectural patterns that must be resolved, lessons learned from other work, and a path forward to an ecosystem of standards-based OVONICA – Open Voice Network Interoperable Conversational Agents. + +ABSTRACT +This is a publication of the Open Voice Network (OVON, www.openvoicenetwork.org), a non profit industry association operating as an open-source community of the Linux Foundation. It asserts that, to realize the full economic and societal potential of conversational assistance, conversational assistants must not only mediate human-to- assistant conversations – e.g., host the conversation and obtain relevant information to fulfill user intents – but also delegate conversations to other assistants, and in doing so, pass textual, acoustic and contextual data, as well as privacy and security controls. +To meet this vision, the Open Voice Network proposes an approach to interoperability between conversational assistants – specifically, the sharing of multi-layered dialogs between assistants and assistants of differing infrastructures through the standardization of communication protocols by which autonomous assistants and assistants collaborate to achieve a common goal. +This paper also suggests areas for further research, as well as next steps for the proposal and testing of communication protocols that may be built from existing technologies and universally accepted standards. +At the core of this work is this firm belief: we are in the early days of conversational assistance, and the future will be one of stunning diversity -- a multiplicity of voice-enabled user end points, content and infrastructure providers, voice-enabled destinations (measured, perhaps, in the billions), industry ecosystems such as transportation and smart homes, and organizational/enterprise value propositions. +This paper is the creation of the Architecture Work Group of the Technical Committee of the Open Voice Network. The Open Voice Network is a trademark of The Linux Foundation. Other trademarks referenced in this report are the property of their respective owners. + + +TABLE OF CONTENTS +1.0 Introduction 6 1.1 Intended Readership 7 1.2 Purpose of this White Paper 7 1.3 The Open Voice Network and Interoperability of Conversational Assistance: Why and What 8 1.4 Going Forward: Your Participation 10 1.5 A User’s Vision: An Interoperable World for Today and Tomorrow 12 1.6 Architectural Aspirations 14 1.7 Boundaries 15 1.8 Interoperability Defined for Conversational Assistance 18 1.9 Core Requirements for Interoperability of Conversational Assistants 19 1.10 Interoperability Intentions of the Open Voice Network 22 +SECTION TWO: A FOUNDATION FOR ANALYSIS -- THE TECHNOLOGY OF CONVERSATIONAL ASSISTANCE 23 +2.0 Introduction 23 2.1 Conversational AI and Voice 23 2.2 Conversational Assistants and Platforms 24 +2.2.1 Conversational Assistants and Agents 25 2.3 Platforms and Content: Existing Standards 26 2.4 Conversational Architecture 26 2.5 The Diversity of Today’s Conversational Assistance 29 +SECTION THREE: AN APPROACH TO VOICE ASSISTANCE INTEROPERABILITY 30 3.0 Introduction 30 3.1 Modeling Interoperability on How People Communicate 30 3.2 Interoperability: Mediation and Delegation 31 3.3 From the Human Model to Standards: OVON Interoperability Objectives 33 + + +3.4 Architectural Patterns Under Investigation 34 3.4.1 Dialog Delegation and Give-Back 336 3.4.2 Delegation Request and Conversation Context 36 3.4.3 Dialog Interaction PayLoad 37 3.4.4 Dialog Component Interfaces 37 +3.4.4.1 An Illustrated Example: Pat Goes Shopping 39 3.4.5 Discovery and Location 40 3.4.6 Sharing and Protection of Data 41 +SECTION FOUR: LESSONS FROM OTHER INTEROPERABILITY INITIATIVES 43 4.0 Introduction 43 4.1 Amazon Voice Interoperability Initiative (VII) 44 4.2 The Stanford Open Voice Assistant Laboratory (OVAL) Model 45 +SECTION FIVE: FURTHER STUDY AND NEXT STEPS 48 SECTION SIX: OPERATIVE VOCABULARY 49 +SECTION SEVEN: ABOUT THE OPEN VOICE NETWORK 53 About The Linux Foundation 54 Acknowledgements 54 +SECTION EIGHT : REFERENCE LIST 556 WORK-IN-PROGRESS APPENDIX 56 +Page 5 of 57 + +PREFACE +The Open Voice Network (OVON) was founded by users of conversational assistance. Its purpose is to develop and drive adoption of technical standards and usage guidelines that will make the emerging world of conversational assistance worthy of user trust. +This document represents the work of but one of several OVON technical work groups – all under the aegis of the Open Voice Network Technical Committee. Separate, yet related initiatives are at present exploring issues of data protection and data security for conversational assistance, assistant-agent destination and location services, voice-centric authentication, and synthetic voice. Other envisioned work streams await resources. +To date, the work has been developed on parallel tracks. As we go forward into Q4 2022, the streams will begin to merge – informing and being informed by each other. +We look forward to collaborating with you. +SECTION ONE: RATIONALE, SCOPE, AND DEFINITION +1.0 Introduction +This section speaks to the Open Voice Network’s reasons for pursuing interoperability of conversational assistants, the boundaries of our work, and a 1.0 definition of conversational assistant interoperability. +Page 6 of 57 + +1.1 Intended Readership +This paper was published for public evaluation and criticism on 12 August 2022. It seeks to address three global audiences, each with a stake in the future of voice: +● Enterprise content creators and communicators (business-to-consumer and business to-business) +● Stakeholders and decision makers within the voice technology industry, including potential architectural partners +● Current and prospective participants and sponsors of the Open Voice Network. +1.2. Purpose of this White Paper +This White Paper seeks to: +● Identify the architectural elements required to facilitate interoperability of conversational assistants +● Prioritize these elements for standardization by the OVON Architecture Work Group ● Foster discussion with the constituents noted above in section 1.1. +In this White Paper, we seek to remain neutral with respect to competing architectural directions, and to reveal the underlying concepts/issues that exist regardless of architectural choice. +This document is intended to be read alongside the Open Voice Network Technical Master Plan. In the short term these two documents may contradict one another as ideas are developed and tested. Such conflicts will be resolved as they emerge. +This paper replaces a previous document entitled ‘Architecture Design.’ +1.3 The Open Voice Network and Interoperability of Conversational Assistance: Why and What +Page 7 of 57 + +The Open Voice Network (www.openvoicenetwork.org) is an open-source community of the Linux Foundation dedicated to the development of the standards and usage guidelines for the emerging world of voice assistance. +It enjoys the regular participation of more than 200 volunteers from 13 nations and 5 continents. The OVON community includes participants from leading voice platforms and infrastructure providers; large enterprises that use voice technologies in customer service and operations (from industries such as retail, healthcare, telecommunications, and financial services); marketing and consulting firms; and more than 40 independent voice development and services companies. +As a technology-neutral, nonprofit organization, the Open Voice Network occupies a unique and strategic position within the voice technology industry. Our sponsors and participants witness a growing diversity of underlying voice technology (speech recognition, language understanding, dialog modeling), conversational design paradigms, labels to represent semantic content, and endpoints (from smart speakers to voice-enabled web pages and immersive games). They also see the rapid growth and evolution of voice value propositions, especially for organizations and across vertical industries. +This diversity points to a significant total available market (TAM) for conversational AI as both an interface and a source for data-fueled insights. Looming before us – as first envisioned by Dr. Monica Lam and colleagues at Stanford University – is an interoperable Worldwide Voice Web (WWVW) (Lam et al., 2021)., where the spoken word offers an open, standards-based interface to billions of voice-enabled media, enterprise, website, transportation, smart IOT environment, and metaverse destinations. +Given the economic and societal value that will be created by an interoperable Worldwide Voice Web, the Open Voice Network commissioned the Architecture Work Group of its Technical Committee to research and recommend architectural options for open, standards based interoperability for conversational assistance. +Using a technologically and architecturally neutral eye, the Architecture Work Group reviewed existing technology standards and existing and proposed voice architectures and messaging protocols. The Work Group ascribed to the Harvard Business Review’s “Jobs to Be Done” methodology (Christensen et al., 2016)., and examined in detail the current and future needs of four key voice constituents: +Page 8 of 57 + +● Consumers of voice-enabled experiences +● Enterprise content providers +● Technical innovators of dialog systems +● Voice system infrastructure providers. +Through these efforts, the work group identified the opportunities and challenges of conversational assistant interoperability, and a series of foundational concepts. +We are deeply grateful for the contributions of the many Open Voice Network participants and sponsors referenced on page 55. These individuals gave freely of their time and intellect to develop this paper and the plans to take it forward. +The Open Voice Network: +● Believes that conversational assistance will realize its economic and societal value when it is interoperable, like telephony or the worldwide web (WWW). +● Sees in today’s rapidly growing enterprise investment (Frost & Sullivan, 2022) in conversational assistance the seeds of a rich ecosystem of independent, purpose- and brand-specific voice assistants, an ecosystem that will bring desired content and experiences to the users of general-purpose, proprietary conversational assistants. +● Recommends the development and industry-wide adoption of standardized communication protocols between conversational assistants operating on different platforms. +● Asserts that conversational assistants must both mediate dialogs (e.g., act as a host, and acquire desired information to fulfill the user intent) and delegate dialogs (e.g., act as an initiator of communication, and bridge to other assistants so that the user intent can be fulfilled.) In addition, conversational assistants must also be identified and authenticated to serve as a destination of a delegating assistant or agent. +● Envisions a model of conversational assistance interoperability in which autonomous assistants and agents collaborate to achieve a goal together. In so doing, they will use standardized ways to share information with each other and operate at various, negotiated levels of trust and information sharing. This will lead to the development, test, and proposal of standards for: +Page 9 of 57 + +o The way in which a spoken assistant name can be used to find an associated assistant +o The way in which basic linguistic information is shared between dialog assistants o The way in which immediate linguistic context and history is shared between dialog assistants +o The way in which control is handed between dialog assistants +o The way in which dialog assistants negotiate trust. +In addition, the Open Voice Network +● Will neither develop nor recommend the standardization of platform components or the format of content used to figure these components. +o We do not believe that a standard methodology for expressing and describing conversational interaction is possible or desirable. Attempts to describe how conversations can or should be modeled will quickly become outdated as conversational interaction innovates and evolves. +In the coming months, the Open Voice Network will develop, test, demonstrate, and propose to existing standards bodies a set of standardized communication protocols that will enable assistant-to- assistant voice interoperability. +1.4 Going Forward: Your Participation +In keeping with the practices of the Linux Foundation, the Open Voice Network works in an open, communal manner. We seek contributions from every corner and region of the conversational AI ecosystem. +Readers of this paper are invited to +● Propose corrections and additions to this paper through the Open Voice Network GitHub repository (https://github.com/open-voice-network) or in the White Papers section of the Open Voice Network website. +(https://openvoicenetwork.org/white_papers). We are especially interested in criticism and comment regarding +o Concepts that are not clearly defined +Page 10 of 57 + +o Requirements that you feel are missing +o Requirements that you feel are unnecessary +● Contribute usage scenarios (similar to those in the paper) that challenge the concepts within this paper, and perhaps point us to new or revised requirements for the interoperability of conversational assistance. +You are also welcome to join the weekly interoperability development meetings of Architecture Work Group of the Open Voice Network’s Technical Committee. One-hour sessions are scheduled Tuesdays at 17:00 CET, 11:00 Eastern, and 08:00 Pacific. Conferencing details are found at https://openvoicenetwork.org/calendar. +Page 11 of 57 + +1.5 A User’s Vision: An Interoperable World for Today and Tomorrow +Today: Pat Shops by Voice as Humans Do +Pat has finished the day’s labors and is driving home. It’s late. A request to Pat’s automobile voice agent starts a music playlist that brings back memories of school days – which, with a start, reminds Pat that a long-time friend from school, now in town for a conference, will be coming to dinner tomorrow. +What to prepare? Perhaps a favorite seafood stew. Crab, mussels, shrimp, snapper in a tomato broth with wine? Pat asks the automobile-based voice agent to connect to the grocer that knows Pat’s preferences – and, in an instant, Pat is directly connected, and speaking to the grocer’s conversational assistant. +A recipe is identified. A shopping list is created -- items, quantities – based upon availability of overnight shipments from seafood suppliers. Wines are recommended (a Zinfandel and Sauvignon Blanc) and chosen. +Pat sighs with relief. And then – right before completed list is transferred to Pat’s mobile grocer app for authentication, pick-up scheduling, and payment -- Pat remembers that the school friend has a chronic medical condition (in this case, diabetes.) Pat quickly asks the grocery assistant if the seafood stew will be healthy or unhealthy. The assistant, in turns, informs Pat that it is directly connecting to a conversational capability developed specifically to recommend food choices for individuals with diabetes. +From the conversational capability, Pat learns quickly that the seafood stew will be a healthy choice for dinner with the school friend. A quick request to the conversational capability, and Pat is sent back to the grocery assistant – and the shopping list is confirmed. +Page 12 of 57 + +Tomorrow: Pat Plans Travel by Voice as Humans Do +Pat plans to travel to an international conference and stay for the weekend that follows. Pat needs a visa, airline reservations, and information about attractions for anticipated free time. Pat’s personal conversational agent is named “Butler,” and was developed by a third party using a mix of proprietary and open technologies. +Pat begins trip planning by pressing a button on a smartphone to begin conversation with Butler. Butler is local to Pat’s phone and has permission to access Pat’s personal data. Butler invokes a voice passport authentication capability to verify that the speaker is Pat and uses Pat’s smartphone for second-factor authentication. +Pat informs Butler of the conference in a major city on another continent. Pat has been asked to deliver a keynote address. A visa may be required to attend. Butler uses a voice destination and location service (similar in purpose to a Domain Name System) to discover the conversational assistant of the passport-visa office of the destination country. +The visa conversational assistant provides Butler with the latest visa requirements and guidance on how (and when) to apply; Butler forwards the guidance and application URL to Pat’s digital account and schedules a reminder to ensure its submission. At the same time, Butler has explored flight options by searching among highly rated travel services. +Butler selects AirWithFlair, a specialty conversational agent and capability for artists and musicians., and mediates a connection to obtain a list of travel options (and prices) for Pat. Butler - knowing that Pat is a loyal customer of the Royalty Hotel chain - delegates the dialog to the Royalty conversational assistant, which recognizes Pat and immediately identifies a location, room type, and pillow choice for the potential conference visit. +Page 13 of 57 + +1.6. Architectural Aspirations +The following architectural aspirations guide our thinking and our actions: +1. We recognize that the material and financial resources required to create, operate, and maintain conversational AI and voice assistant systems are significant. Our architectural direction will allow for sharing the cost of those resources, but we will not ourselves specify how that needs to occur. +2. There are several competing architectures for control interfaces between conversational AI platforms, assistants, and agents. OVON will evaluate architectures, produce recommendations, and attempt to remain neutral with respect to the choice until and unless making a choice becomes required for completing our specifications. +3. Each of the dominant proprietary platforms within the existing voice market has an installed base of platform-specific applications. At the same time an ecosystem of third party voice content creators has emerged outside the proprietary platforms. To maximize the value of our work and speed global adoption, we will seek to establish a platform-neutral approach to interoperability – while adopting (as noted above) common concepts and ideas. +4. The lifespan of a standard is much longer than the lifespan of the physical components subject to the standard. The specifications and guidance which we will produce may be informed by existing design patterns but must not be limited to those patterns. +5. In the immediate, we aspire to: +a. Outline the different architectural approaches that could support +interoperability. +b. Outline the advantages and disadvantages for each with respect to the OVON constituents and the Jobs to Be Done. +c. Recognize that several viable architectural approaches to interoperability may coexist and that OVON needs to remain responsive to the directions taken by the market. +d. Identify platform-neutral standardization that will enable both the new emerging third-party ecosystem and the applications resident on proprietary voice platforms. +Page 14 of 57 + +When you’re exploring the internet, do you shut your browser down to open a new web page? +Why, then, would you want to do that with voice? +Royal O’Brien, Linux Foundation, July 2022 +1.7 Boundaries +The Open Voice Network defines the boundaries of this work using the “Jobs to Be Done” approach (Christensen et al., 2016)., singling out four groups of individuals for whom interoperability will create incremental and sustainable value. Identification and clarification of the needs and desired experiences of each group allows the OVON to define the work required from the Open Voice Network to fulfill those needs and desires. +Though Open Voice Network-delivered standards will benefit the voice industry at large, our work suggests that four distinct groups of individuals are central to the question of interoperability: (1) consumers of enterprise content who wish to use voice enabled technologies, (2) enterprises who provide content and wish to provide voice enabled experiences, (3) technical innovators of dialog systems who wish to create the systems for the enterprises to use, and (4) voice systems infrastructure providers who wish to create and pursue commercial opportunities within a free and open environment. +Page 15 of 57 + + +Figure 1.7 Benefactors of conversational assistant interoperability; the value ecosystem +Together, these groups constitute a tiered ecosystem that will flourish in a world of conversational assistant interoperability. +Consumers of Voice Enabled Experiences are at the top of the tier. These are the end consumers of voice enabled experiences. OVON-developed standards will enable them to engage in voice experiences in web environments that currently do not provide voice experiences. The “Jobs to be Done” that we are supporting for this group include the following tasks: +● Help me continue my voice task as I switch voice devices (home, phone, auto) ● Help me avoid restricted access to data in all internet-connected, voice-enabled systems ● Help me opt-out of data protection when it makes sense. +● Help me avoid restrictions on vendor choice. +Enterprise Content Providers rank next. These are the businesses and organizations that offer applications on proprietary, general-purpose consumer platforms, or do not deliver voice experiences to their customers outside the organizational security firewall. Together, enterprise content providers represent a significant third-party content and voice-centric services ecosystem. “Jobs to be Done” for enterprise content providers are: +● Help me avoid building separate applications and services in order to use all global voice platforms and web / IoT / metaverse-based voice-based destinations +● Help me grant all potential constituents free and unfettered access regardless of the +Page 16 of 57 + +constituent's "home" voice platform +● Help me ensure that all my potential constituents can freely and directly connect with my voice applications and services, regardless of home platform +● Help me enable the privacy of all dialog data (text, acoustic, semantic) shared between my applications and services and my constituent when explicitly requested to do so (however that may be) +Technical Innovators of Dialogue Systems are the technology-centric and consulting organizations that will create customer service, web, or operational environments for enterprise content providers. These innovators develop the different technology solutions that work together to enable voice experiences. They will build the pieces of the modularized future we envision. “Jobs to be Done” for technical innovators include: +● Help me avoid creating new voice-based core business processes for transactions ● Help me avoid creating new/redundant messaging protocols for sharing of dialog, data, and controls between platforms, agents, and assistants +● Help me give customers the choice to protect their voice data +● Help me secure my customer's voice data +● Help me avoid creating new specifications /processes for voice authentication, attestation, and authorization +● Help me get to content and consumers without dependence upon Big Tech proprietary platforms +Of note: while there are “Jobs to Done” for each of the constituent groups – and each plays an important role in an interoperable ecosystem – all will rely on the work of the Technical Innovators. These individuals will use Open Voice Network standards to do their work; the OVON expects to provide significant support to assist them in accomplishing their jobs to be done, each of which require a specific experience. These are the experiences we will create for the technical innovators. +Technical Innovator Jobs to Be Done Corresponding Technical Innovator Experiences +Help me continue my voice task as I switch voice devices (home, phone, auto) Implement OVON standard protocols and specifications to enable business processes for interaction +Help me avoid creating new/redundant messaging protocols for sharing of dialog, data, and controls between platforms, agents, and Select and plug and play messaging protocols + + +Page 17 of 57 + +assistants +Help me give customers the choice to protect their voice data My customer elects to protect their voice data +Help me secure my customer's voice data I match security level per customer's level of data protection +Help me avoid creating new specifications /processes for voice authentication, attestation, and authorization I learn OVON Standards so that I know how it should work +Help me get to content and consumers without relying on the proprietary Big Tech I know where to go and what to do to manage proprietary Big Tech platforms + + +Voice Systems Infrastructure Providers provide technologies used by content creators and technical innovators to deliver voice-based content and services to consumers of voice-based experiences. In an open, standards-based, interoperable world, they will enable a stable voice ecosystem. Technical innovators will rely on these open infrastructure offerings to create voice technology solutions that will be used by enterprise content providers. “Jobs to be Done” that we are supporting for voice system infrastructure providers are the following: +● Help me create an accessible environment +● Help me create an environment where seamless interaction is commonplace ● Help me create an environment where scalability is empowered +Two additional jobs to be done were identified during working group consultation with a major infrastructure provider: +● Help me maintain ownership of data +● Make sure that when my infrastructure is used (work is being done by me), that my work gets recognized. +1.8. Interoperability Defined for Conversational Assistance +At the most basic level, interoperability means “the ability of two or more systems or components to exchange information and to use the information that has been exchanged.” IEEE Standards Information Network/IEEE Press (2000). The OVON adapts this definition to the conversational assistance ecosystem as follows: +Page 18 of 57 + +Interoperability is 1) the ability of conversational assistants of diverse parentage to collaborate seamlessly to achieve a goal together, using standard ways to share information with each other, operating at various levels of trust and information sharing, and operating independently of the physical devices being used; and 2) the ability of conversational AI-driven content, assistants, and service providers to reach target audiences across the diversity of platforms, general purpose assistants, and physical access devices (endpoints). +Conversational assistants can be fully or partially interoperable. +1.9 Core Requirements for Interoperability of Conversational Assistants Conversational assistants must: +● Be enabled to fulfill user intents through mediated and delegated communication. ○ In mediated assistance, the user interacts with a single conversational assistant. The mediating conversational assistant hosts the conversation; it never cedes control of the conversation. To fulfill user intent, the mediating assistant may obtain information from third party sources or introduce an application resident on its platform. +○ In delegated assistance, the user interacts with one or more independent conversational assistants using standardized messaging protocols. The hosting conversational assistant initiates the conversation; it passes control to a delegate assistant (or assistants). +● Be enabled to serve three roles to fulfill user intents: as 1) a host of applications that mediates content to fulfill user intent, 2) an initiator of a connected dialogue that delegates fulfillment of intent to one or more independent assistants, or 3) a destination that is delegated to, and as such, assists the initiator in fulfilling user intent. +Page 19 of 57 + + Figure 1.9: Core Requirements for Conversational Assistants +As an initiator and as a destination, an assistant must offer six types of interoperability, as shown in the grid below. +Type of +Interoperability Use Example Benefit +Share dialog +fragments +Dialog fragment = portion of a dialog pattern, a piece of interactive +fragment, a chunk of interaction +specification Conversational +assistant uses +fragments from +one or more other conversational +assistants to +achieve a dialog goal – user thinks is talking with one assistant A delivery service conversational assistant (with Pat’s consent) uses the location collection dialog +fragment from a payment +conversational assistant to collect Pat’s address information. The fragment may also be identified so it can be reused. Application +developers save time by reusing +dialog fragments from other +conversational +assistants. + + +Page 20 of 57 + +Share personal data +Personal data = information which the user wants to maintain across several sessions. +A conversational assistant specifies how data is shared with another +conversational assistant. +The value in the Amount Due slot in the Shop conversational assistant is copied into the Payment Amount slot in the Payperson conversational assistant. +User controls how personal data can be shared with other +conversational assistants, saving users from +reentering the data value. + +Share +conversational context +Conversational +context = +information which is shared between a user and a +conversational +assistant over a (tbd) period of time. With the user’s permission, +conversational +context may be shared with other conversational +assistants. Relevant context (of the user, of the dialog) established on one assistant is shared with +another +conversational +assistant. + The context (history) of actions in the Shop voice assistant is shared with the Payperson voice assistant. The Payperson conversational assistant is aware of information collected and generated by the Shop voice assistant. Enables +conversational +continuity +conversational +continuity = the +principle of making sure that all details in one voice +assistant are +consistent with the details in another voice assistant +Transfer control +Transfer control = determine which conversational +assistant is invoked and becomes +active. +Determination of which assistant has “the floor,” which assistant has the Transfer control from one +conversational +assistant to +another While shopping for groceries +through a conversational assistant, Pat switches to the Payperson voice assistant to pay for the groceries. Matches the +natural thought +processes of +human users. + + +Page 21 of 57 + +permission to +speak. +Share +components (This is an aspirational goal) +Components = +components of a conversational +assistant, +including but not limited to ASR, +TTS, NLP, NLG. The developer +specifies which +components are used within a +dialog. A voice developer may desire to use sophisticated natural language processing rather than a simple natural language processor provided by the conversational assistant being used. +Other examples of component sharing include (a) using different ASRs for speakers of American English and British English, (b) performing natural language +translation between national languages such as English/German or English/Mandarin, (c) providing access to algorithms that perform explicit query or implicit query (a.k.a. voice search). Developers can +develop custom assistants through the use of +standards-based, best-of-breed +technologies. +Share endpoints +Endpoint = +hardware or +software through which the user +speaks and listens to conversational assistants User controls +where input may come from or +output may be +directed to. The user may begin a request from a device with a public +microphone/speaker and later request to continue with a private device. +The user may ask for output to be directed to a different +conversational assistant or a +different device, such as a printer, or may ask for input to come from some shared data. User can access +any conversational assistant using any endpoint + + +Page 22 of 57 + +1.10 Interoperability Intentions of the Open Voice Network +The Open Voice Network seeks to enable mediated and delegated communication by conversational assistants within an ecosystem of general purpose, consumer-centric conversational assistants, and independent, purpose- and brand-specific conversational assistants and assistants. To make this happen, OVON will develop messaging protocols and a destination management and services system, to standardize the way in which: +● a spoken assistant name can be used to find an associated assistant +● basic linguistic information is shared between dialog assistants +● immediate linguistic context and history is shared between dialog assistants ● control is handed between dialog assistants +● dialog assistants negotiate trust. +SECTION TWO: A FOUNDATION FOR ANALYSIS -- THE TECHNOLOGY OF CONVERSATIONAL ASSISTANCE +2.0 Introduction +The Open Voice Network grounds its research into conversational assistant interoperability in current operative definitions of conversational Artificial Intelligence, as well as the normative components and architecture of contemporary voice assistance. In this section, intended for +both subject matter expert and lay audiences, we share our understanding of those components and architectures. +2.1 Conversational AI and Voice +Voice assistance is but one form of conversational Artificial Intelligence, or conversational AI. +Page 23 of 57 + +Citing Interaction.com (2022): “Conversational AI is the set of technologies behind automated messaging and speech-enabled applications that offer human-like interactions between computers and humans. +“Conversational AI can communicate like a human by recognizing speech and text, understanding intent, deciphering different languages, and responding in a way that mimics human conversation.” (Haas, 2020). +Conversational AI solutions can be offered through text and voice modalities. Voice solutions – referred here broadly as voice assistance – can be delivered as an audio-only interface, or one in which voice assistance works in a complementary way with screen-based information. The latter is often described as multi-modal voice assistance. +This paper speaks primarily to the sharing of voice-based elements within a dialog. However, the Open Voice Network recognizes clearly the importance of multi-modal voice assistance (especially in an enterprise context) and will extend future editions of this work to multi-modal interoperability. +2.2. Conversational Assistants and Platforms +The Open Voice Network draws a clear distinction between conversational assistants and conversational platforms. +Conversational assistants function as conversational user interfaces. They engage in conversation and are identifiable by users as conversational actors. A conversational assistant typically has an assigned identity, an identifiable voice, and/or a name: for instance, Alexa, Siri, or Magenta assistants. The user recognizes a given conversational assistant as a singular entity and refers to that assistant with a pronoun (he, she, it, they) or by an assigned name. The conversational assistant may also refer to itself using a name or a pronoun. +Conversational assistants often respond to queries in the form of web searches. The response typically describes how the assistant locates the delivered content. For instance, one might ask, “Assistant, what is the weather like in Brooklyn, NY today?” To which the assistant responds, “Okay, here is something I found on weather.com . . . “ +A conversational platform is the software to which the assistant connects, and that interprets the input from the user. The assistant responds through an endpoint, which is typically a piece of hardware, such as a smart home display, a television remote control, an automobile smart dashboard, a smartphone, or a smart speaker. +Page 24 of 57 + +There are endpoints and platforms currently associated with proprietary interfaces, such as Amazon Alexa or Google Assistant. They handle wake words (“Hey, Google…!”), audio coding and streaming, and graphical events and content. +If a given platform provides generic access to content, it may have a host voice assistant, and give the user the option to ask for a delegate assistant. In other words, the host can provide a gateway to delegate content. +With the exception of assistants written in the VoiceXML standard (McGlashan et al., 2004)., delegate assistants are typically authored for a given platform by third party companies. +Examples of End Point-Platform-Assistant Pairings +End Point Platform Host Assistant +Echo Dot Amazon Alexa +Google Home Google (Hey) Google +Magenta Mini Deutsche Telekom Magenta + + +Names and brands may be the property of others. +2.2.1 Conversational Assistants and Agents +For the purposes of this paper, the terms “assistant” and “agent” both refer (as noted above) to conversational AI user interfaces that are perceived to be a single conversational actor, operate on behalf of the user, and operate in accord with a conversational platform. Both assistants and agents have an addressable name, continuity of knowledge, and a bounded identity. +Both assistants and agents have the ability to fulfill user intents, mediate and delegate dialogs, and serve as a destination. Agents will also operate independently on behalf of the user. In the worldwide voice web, there will be specific-purpose conversational assistants and agents, and general-purpose conversational agents. specific-purpose conversational assistants and agents, and general-purpose conversational agents (McGlashan et al., 2004). +Page 25 of 57 + +2.3. Platforms and Content: Existing Standards +Most conversational systems separate the platform from the content that is executed on that platform. In Web Models - a common structure - the content itself is typically either static or active web content. The platform is roughly analogous to a browser; the content is stored on web servers and accessed by standard web protocols (e.g. HTTP) (Fielding, 2022). +A set of related standards proposed by the W3C Voice Browser group has successfully created standardized formats for describing content for spoken dialog, speech recognition grammar, semantics, text-to-speech markup and pronunciation (McGlashan et al., 2004). This content can be interpreted by a W3C compliant voice browser. +The W3C voice-browser group standards have served the telephony community well but have not yet been adopted by the leading general-purpose conversational platforms, nor by the many creators of text-centric chatbot platforms. One notable exception is speech-synthesis markup, which has been widely adopted (McGlashan et al., 2004). +This new wave of voice assistant platforms also uses a web-content model to support third party dialog content. The formats for such content are mostly published but remain proprietary. Examples include Amazon Skill Definitions and Google Actions. Content for this new wave of platforms utilizes standard web protocols such as HTTP or OAuth (Fielding, 2022), but also strongly promotes or mandates proprietary infrastructure and authoring tools. +2.4 Conversational Architecture +Below is a schematic view of a generic conversational assistant system. +The Open Voice Network bases its explorations of voice assistance interoperability upon this understanding of system components and architecture. +Page 26 of 57 + + Figure 2.4 Typical Components of a Conversational Assistant +Figure 2.4 illustrates the schema of traditional conversational architecture, along with the various components that are collectively enlisted to enable two-way communication with the user. +The architecture typically operates as follows, as per the above diagram, from left to right: +The user begins (far left) with an endpoint, which can be located on any piece of digital hardware. Today’s endpoint options include conversational assistant implementations on devices ranging from smart speakers, smartphones, desktop or laptop computers, and smart home systems, to automobile dashboards, television remote controls, and kitchen appliances – the list is endless. Endpoints provide a software or hardware interface between the user(s) and the Conversational Platform, which includes all the components within the encircled figure above. An example of a simple endpoint is an old-fashioned telephone (often referred to as the +Page 27 of 57 + +plain old telephone system, or POTS). A sophisticated endpoint might be an immersive VR environment.1 +Endpoints engage with the user via audio, text and visual media. Platforms interpret content that embodies voice assistants and enables them to interact with users via the endpoint. +The diagram above depicts three conversational components for exchanging information between the endpoint and the interaction manager. Let’s explore these as listed from top to bottom in Figure 2.4. +The most traditional form of user-assistant communication involves Speech Recognition Technology, where the device interprets the user has uttered. Speech recognition technology (speech-to-text) has grown more accurate and sophisticated in recent years. Previous generations of this technology enlisted a very narrow spectrum of comprehension; at each dialog step, they required application-specific grammar to define the wording of 'allowable' spoken responses. Such grammars could be described using the W3C GRXML standard (McGlashan et al., 2004). Newer generations of speech-to-text engines use general purpose language models that can be shared across applications and contexts. This, of course, represents a considerable step forward. However, there are still no open standards for language model description at the present time, which limits interoperability. +Empowering this process, Natural Language Understanding (NLU) capability - a branch of Artificial Intelligence - enables the comprehension of the meaning of a user’s human speech and clarification of underlying intentions. +The second category of conversational components - a Text-to-Speech System (TTS) - serves to convert language text into audible speech and can be implemented in software or hardware products. Normal language text may contain symbols, numbers and abbreviations that are converted to spoken words presented to the user via a speaker. The quality of a speech synthesizer is gauged by its similarity to the human voice and by its ability to be understood clearly. +The third category of conversational components, the Graphic User Interface (GUI) Adapter, comes into play in those unique instances where the user requests rich media content, such as buttons, images, and formatted text, to be visually displayed on the endpoint’s screen. The +1Standards exist for establishing audio connections between endpoints and platforms, for example in POTS telephone systems or VoIP connections. Smart speaker endpoints currently use interfaces that are specific to the platform. These may be published but they are not interoperable with other platforms. +Page 28 of 57 + +endpoint will generally only be able to display said content if it receives specially formatted input from the platform. Likewise, the platform will only be able to “understand” user GUI input if it is first “translated” into a form that platform can understand. +As defined by the World Wide Web Consortium (W3C) Voice Interaction Community Group (McGlashan et al., 2004), Dialog Manager is a component that receives semantic information derived from user input (via speech recognition + NLU), updates the dialog history, its internal state, then decides upon subsequent steps to continue a dialog and provides output. +The Dialog Manager gets its content and logic from the Web Content layer. The content might be hard coded or be dynamic via interactions with external APIs (for example it could output an address using a third-party address normalization API). +The Dialog Manager also passes information to the Session Manager and retrieves information from it. The Session Manager is a governance body within the voice architecture. It governs and administers conversational sessions with the user. This means controlling starting and stopping +points, identifying and verifying the speaker, verifying the conversational assistant, managing connectivity among multiple conversational assistants, and, significantly, maintaining the context necessary for an active conversation between a user and a bot/assistant. A unique session ID is generated for each session, allowing the context of the conversation to be maintained over time. +2.5 The Diversity of Today’s Conversational Assistance +The popularity of general-purpose consumer conversational assistance (Amazon Alexa, Google Home, Deutsche Telekom, Samsung Bixby, Baidu Xiaodu, and others) (Frost & Sullivan, 2022) has often obscured the growth of other parts of the conversational assistance ecosystem. +Firms such as Microsoft, Deutsche Telekom, Nuance (now a division of Microsoft), SoundHound, PeopleReign, Cerence, RASA, Mycroft.ai, Redfox.ai (and many others) (Frost & Sullivan, 2022) now offer enterprises the tools and services to develop conversational +Page 29 of 57 + +assistance solutions for devices, smart environments, personal transportation, customer service and employee support. +Enterprise investment in conversational AI worldwide is expected, by 2025, to reach a level double that of where it was four years prior; we’re now seeing per year annual growth rates of roughly 22-25 percent (Frost & Sullivan, 2022). +Analysts foresee the bulk of future growth in these areas: +● Customer service and employee support solutions – many of which have grown through the years from call center automation. These are often classified as interactive voice response (IVR) systems. +● Personal and autonomous transportation +● Smart environments, from homes to manufacturing facilities +● Hands-free environments, from surgical suites to warehouses + +SECTION THREE: AN APPROACH TO VOICE ASSISTANCE INTEROPERABILITY +3.0 Introduction +Section Three describes the Open Voice Network approach to the interoperability of voice assistants, and the current intent fulfillment patterns under investigation. +3.1. Modeling Interoperability on How People Communicate +The Open Voice Network acknowledges that a standard methodology for expressing and describing conversational interaction may be achievable -- but may not achieve the objectives of the ‘Jobs to Be Done’’ identified above, particularly for the Technical Innovators. +Attempts to describe (and standardize) how conversation can and should be modeled will quickly become outdated as the field of conversational interaction innovates. As an example, +Page 30 of 57 + +the latest generation of Voice System Infrastructure Providers and Technical Innovators put aside the widely adopted W3C VoiceXML and Speech Recognition Grammar standards in favor of creating new design paradigms based on less constrained speech-to-text and NLP technologies. +We assert this: human languages can be considered the ultimate interoperability standard. People have internal hidden motivations and knowledge. They use language to align their understanding of the world with tasks to be achieved. In some conversations people will have enormous levels of trust and shared understanding prior to the interaction. In others they will be complete strangers and may not trust each other. +We envisage that successful interoperability approaches will support a model of autonomous agents collaborating to achieve a goal together using standard ways to share information with each other, operating at various levels of trust and information sharing. +3.2 Interoperability: Mediation and Delegation +The modeling of voice assistant interoperability according to human communication leads us to identify two types of interaction: mediated and delegated. + Figure 3.2.a. Example of mediated and delegated communication. +Page 31 of 57 + +Figure 3.2.a shows two examples of how people inter-operate to achieve a conversational task, in this case between a 'user' and a helping ' assistant'. +In the top example of mediated communication, the user interacts with a single assistant. assistant A owns the conversation but may talk with others (e.g., assistant B) behind the scenes to achieve the goal. The client may be aware of the existence of assistant B but does not interact directly with them. Assistant A is the host of the communication. +The second example is one of delegated communication. To fulfill the intent of the user, assistant A passes control of the conversation to assistant B. Assistant A may monitor the ongoing interaction or may drop-out of the conversation altogether. In this example, assistant A is the initiator of the communication. +In this example of delegated communication, assistant A may need to talk behind the scenes with assistant B to establish whether assistant B is indeed the appropriate (or best) destination for delegation prior to their actual act of delegation. +To properly fulfill all user intents, assistant A will need to both mediate and delegate conversations. In some cases, the client may not need to communicate with assistant B resulting in a fully mediated conversation. For example, assistant A may place the user ‘on hold’ if this were a telephone conversation, or the dialog may be paused if it is an electronic channel such as chat. +In addition, assistant A may also be asked to serve as a delegate destination – to receive a request from assistant C to fulfill a user intent. +Analogous to this, Figure 3.2.b illustrates another aspect of interoperability between people. In situations where considerable trust exists between the collaborating assistants then brief referrals (in language) can be made along with the passing of a case-history. +Page 32 of 57 + + +A brief referral plus a detailed case history is passed to the next specialist. Each specialist interprets their role based on a referral request and the case history +Figure 3.2.b. Example of a delegated communication in which context (here, a cast history) is also shared. +The Open Voice Network proposes that, to the extent it is possible, models of inter-assistant and inter- assistant collaboration follow similar patterns to those followed by human assistants. +Key features of such a model would include: +● Each assistant has an identifiable identity and knowledge boundary. +● Each assistant will both mediate a communication (serving as a host assistant) and delegate a communication (serving as an initiating assistant) depending upon the user intent. assistants will also serve as a destination of a delegated communication. +● A desired scheme for interoperability will support collaboration as a mixture of mediated and delegated patterns. +● Delegation between assistants will be performed using brief requests. +● Levels of trust will vary between clients and assistants and this will affect the level of information sharing across boundaries. +● Where high levels of trust exist and there is shared understanding of how to represent knowledge, delegations and mediations will be supported by shared history and context. +Page 33 of 57 + +3.3 From the Human Model to Standards: OVON Interoperability Objectives +The OVON architecture seeks to establish standardized communication protocols between dialog assistants running on diverse platforms. +It does not, however, seek to standardize platform components or the format of content used to configure such components. OVON envisages a world in which conversational systems will continue to evolve in complexity and function. +As such the OVON architecture celebrates diversity in the following areas: +● The diversity of underlying technology (i.e., speech recognition, language understanding and dialog modeling) +● The diversity of endpoints [e.g., smart speakers, smart phone apps, dumb phones (POTS), audio-enabled web pages etc. +● The diversity of conversational design paradigms +● The diversity of labels used to represent semantic content. +To enable mediated and delegated communication by conversational assistants within an ecosystem of general purpose, consumer-centric conversational assistants, and independent, purpose- and brand-specific conversational assistants and assistants, the Open Voice Network will specify the messaging protocols and a destination and location services system that will provide a standardized approach to the way in which: +● basic linguistic information is shared between dialog assistants +● immediate linguistic context and history is shared between dialog assistants ● control is handed between dialog assistants +● dialog assistants negotiate trust. +● a spoken assistant name can be used to find an associated assistant. +Page 34 of 57 + +3.4 Architectural Patterns Under Investigation +In response to the Open Voice Network’s Jobs To Be Done review, and informed by the vision above, the Open Voice Network Architecture Work Group identifies several architectural patterns that, when combined together, could enable seamless intent fulfillment for users of different conversational assistants and assistants. +The architectural patterns that are being explored are: +● Dialog Delegation and Give-Back (additional detail to follow in Section 3.4.1) +A standard, distributed approach to the delegation and take-back of tasks from one platform to another. This is modeled on the idea of intercommunicating human assistants. +● Delegation Request and Conversation Context (Section 3.4.2) +A standard, extensible method of passing context between assistants. This is closely related to the delegation and take-back mechanism. This is modeled on the idea of maintaining case-history during sustained interactions. +● Dialog Interaction Payload (Section 3.4.3) +A standard, extensible format for the passing of linguistic events between components. This is expected to consolidate work done over many years by the linguistic and dialog engineering community. +● Dialog Component Interfaces (Section 3.4.4) +Standard component interfaces, built on top of the linguistic event format, starting with dialog interaction managers. This is expected to align as closely as possible with existing proprietary approaches. +● Discovery and Location (Section 3.4.5) +Standard patterns and mechanisms to allow assistants to publish and discover services from one another. +● Sharing and Protection of Data (Section 3.4.6) +A strategy for sharing data elements among conversational assistants. It also postulates a constraint enforcement mechanism that enforces privacy +constraints on shared data. +Page 35 of 57 + +These six different design patterns could be used separately or in conjunction with each other. They could also be adopted as part of other interoperability initiatives such as the Stanford University Open Voice Assistant Lab (OVAL) model discussed later in this paper. +Additional design patterns are in review by the Open Voice Network Architecture Work Group. Initial work on those patterns can be found in the Appendix, Section 1. +3.4.1. Dialog Delegation and Give-Back + +Figure 3.4.1 Delegation request +Figure 3.4.1 illustrates a typical delegation request. The user’s request is transferred from the user to a conversational assistant (which we will call the host) which examines the request. If the host determines that it cannot process the request itself, then the host negotiates (the dotted line in figure 3.4.1) with a second conversational assistant (the delegate) for possible processing. +A high-level acceptance/rejection mechanism delegates the request to the second conversational assistant (delegate conversational assistant). Some establishing or authenticating information will typically accompany the request to provide context about the initiating assistant and requested task. +Through this mechanism, the delegate conversational assistant will assume control, and then fulfill the user’s intent. In the most basic version of this scenario, the delegate assistant will then declare the task as finished. It may also ask the host assistant to reassume control. +Page 36 of 57 + +3.4.2 Delegation Request and Conversation Context +When, to fulfill a user intent, an initiating assistant seeks to delegate the conversation, the initiating assistant will frame its request so that all potential destination assistants can understand it. An establishing context for the dialog may also be necessary. +A delegation request from an initiating assistant requires these capabilities: +● Request: what the host assistant would like the potential delegate assistant to do: the user request +● Context: conversational history, plus other parameters (TBD). +3.4.3 Dialog Interaction Payload +The request payload contains information in standard formats that OVON will define. It consists of: ● Current request +● Conversational History +● Other data +The process of delegation may begin with an initiating assistant ( assistant X) passing part of all of the user utterance as its ‘request’ to a Destination and Location Service or directly to a destination assistant ( assistant Y). For example, one might say, “assistant X…” (wake word) “. . . “. . send me to assistant Y . .” ‘. . so that I might accomplish x, y and z”. +The potential destination assistant Y will determine if and how it can respond to the request. +The Architecture Working Group is exploring the use of a standardized dialog ‘payload’ format to support the passing of requests in a standard form. +3.4.4 Dialog Component Interfaces +The OVON Architecture Working Group identifies three levels that can be used to communicate a request from one assistant to another. +Page 37 of 57 + +● Level 0. An audio stream or file containing a natural language request. ● Level 1. A text string containing a natural language request +● Level 2. A structured semantic representation of a natural language request. +There is no architectural requirement for the delegation request to be built directly from the user’s language. The host delegate could synthesize a delegation request in its own natural language. When level 2 communication is used, both conversational assistants must have a common understanding of the semantic representations. In addition, we do not anticipate a requirement for a delegation request to be built directly from the initiating assistant’s language. The initiating delegate could synthesize a delegation request in its own language – natural or internal. This would also be of value in situations in which multiple natural languages are present, with either a multilingual human user or a multilingual assistant. The mechanisms for handling multilingual delegation are not yet specified. +In addition to the three layers noted above (audio, text, and semantic representations of speech) the following may also be represented in a conversational history: +● a user ID, standard for Identifying users +● a platform ID, standard for identifying the platform +● a conversational assistant, standard for identifying the conversational assistant ● a session ID, standard for identifying session +● inflection, standard symbol sets for annotating stress and intonation of speech ● affect, standard symbol sets for coding of emotions or affect of the speaker ● a dialog act, standard symbol sets for representing dialog acts (e.g., asking a question). +Other richer semantic representation formats will likely emerge over time. The set of supported layers will grow over time as proprietary formats are proposed for adoption. +The Open Voice Network may propose a standard format for each of these layers. At present, we envision the identification of specific layers as either mandatory or optional. +Regarding the Open Voice Network approach to standardization, the development of each of these layer schemas will draw heavily from formalized or de facto industry standards. These could include formal standards (e.g., TEI, UTF, W3C Pronunciation representation, etc.) or broadly adopted specifications from industry leaders. +Page 38 of 57 + +Each layer may possibly also be expressed using the Extended MultiModal Annotation Language EMMA (Johnson, 2009). +In an early proof of concept demonstration, the OVON Architecture Work Group successfully transmitted level 1 (text) messages between Mycroft, Magenta, and Genie conversational assistants. [MOAD summary.docx - Google Docs]. +3.4.4.1 An Illustrated Example: Pat Goes Shopping +Let’s return to the illustration of shopping with Pat, as presented earlier in section 1.5. Below, figure 3.3.4 breaks down the requests of Pat’s dialog with two conversational assistants, Shop (a home shopping assistant) and Payperson (a payment assistant). Note the three layers of interfaces and their interaction, as depicted here. + Figure 3.3.4: Current request and a subset of the conversational history of a Pat shopping dialog. +● Layer 0: Audio. An audio stream or file containing a natural language request. Voice is represented in computers as a sequence of bits, often called a wav file. Above, we show electronic representations as wave icons. +Page 39 of 57 + +● Layer 1: Text. A text string containing a natural language request. Above, the red icon, representing Pat’s utterance, is converted into the text: “Put two apples in cart.” +● Layer 2: The structured semantic representation of a natural language request, here shortened to "goals." Semantics (meaning) of a user goal is often represented using intent and slot container structures, which represent goals/actions to be performed and the parameters defining the actions. Above, Pat’s text expression is converted into the Intent, PutInCart, with the parameters “Apples” as the value of Food Item and “1” as the value of Quantity. +To complete the example, the dialog manager generates the text message, “Two apples were placed in cart,” which is converted by the TTS into a wav file (depicted by the green waveform graphic above) and then presented to the user. +. +3.4.5. Discovery and Location +An open, worldwide voice web will allow users to fulfill an intent through a mediated or delegated connection to any content source or conversational assistant, regardless of platform parentage. We believe that conversational assistance can and should work like the web. +Users will need to both find (e.g., the connection to a specific, named destination) and discover (e.g., the exploration of information and destination options within a topic or category) content in the worldwide voice web. +To meet this need, the Open Voice Network has initiated research into the development of what may be termed Discovery and Location Services. The services may include: +● a DNS-like service for the identification of available conversational assistants (by name and address) +● A standardized approach to metadata representation and organization +● the use of search engines, browsers, and aggregators that examine meta information about conversational assistants that provide the conversational assistant’s address +● the use of natural language requests sent from one assistant to another which the receiving assistant can use to decide if it can fulfill the request. +Page 40 of 57 + +A request to a Discovery and Location service from an initiating assistant will likely include text with the user’s request and contextual information. +The response from Discovery and Location services(s) and an identified destination assistant will include a digital address, configuration information, and privacy/security constraints. + +Figure 3.3.5: Schematic of the role of an envisioned Discovery and Location Service. +The schematic above illustrates the role of an envisioned Discovery and Location Service in the process of delegation. In a future worldwide voice web with millions, if not billions of destinations, a Discovery and Location Service (or Services) would enable an initiating assistant to find and connect with a destination assistant. In practice, it would determine which conversational assistants should be invited to the delegation and identify the digital addresses of the potential destinations. +3.4.6. Sharing and Protection of Data +An open, standards-based worldwide voice web of mediated and delegated dialogs will require a new and systemic approach to data protection. The Open Voice Network seeks to enable secure end-to-end connection between users and conversational assistants. This may include the encryption of messages, authentication of users as they initiate and progress through delegated dialogs, and authentication of conversational assistants in both mediated and delegated dialogs. +Page 41 of 57 + +Work in these areas is underway. +Earlier this year, the Privacy and Security Work Group of the Open Voice Network Technical Committee published for public review the guiding white paper Privacy Principles and Capabilities Unique to Voice. +From a perspective of a human user of voice assistance – and anticipating an OVON assertion of both mediated and delegated interoperability – the Privacy Principles white paper reviews the current landscape of regional and national privacy regulation and legislation, identifies primary human user risks and potential harms, and details four principles that the OVON believes should guide the development and implementation of voice assistance: +● Transparency: proactive communication – through an easily accessible and readily available interface -- to the individual user of data collection practices, general data usage, and data sharing policies. +● Consent: the expectation that the individual user must be allowed to give explicit, unambiguous agreement to the collection and processing of personal data. +● Limited Collection and Use: a restriction of the collection and analysis of raw and processed voice data – beyond that necessary for immediate dialog functionality – to stated and creator consent-given purpose. +● Control: the ability of the individual user to easily access, rectify, suppress, limit, oppose, and transport the data created by the individual user. +Under development now by the Privacy and Security Work Group are two additional papers: +● Data Security Specific to Voice, a review of current data security issues within voice assistance, and of data security issues inherent within multi- assistant voice interoperability. This paper is scheduled for publication in August 2022. +● Voice Assistance and the Privacy of Organizational Data, an exploration of why, when, and how organizations of all types (and especially enterprises) may protect proprietary data from undesired sharing and platform collection in both mediated and delegated interoperability. A first draft of this paper is scheduled for publication in the fourth calendar quarter of 2022. +Page 42 of 57 + +In parallel, the Ethical Use Task Force of the Open Voice Network published for public review the white paper Ethical Guidelines for Voice Experiences. +From the perspective of a human user of voice assistance, the Ethical Guidelines white paper identified voice-specific issues of ethical and moral concern, discussed rights that must be respected and values to be promoted, and addressed preventative measures that could be taken to protect users from voice-specific harm. +The paper also outlines a voice-specific ethical framework of five principles: +● Compliance: the acknowledgement of, and adherence to, ethical principles, standards, guidelines, and existing laws and regulation. +● Transparency: open and clear communication – in an easily accessible, understandable, and explainable user interface – regarding the collection, usage, and sharing of user and user-created data. +● Privacy Protection: not only adherence to the regulation and legislation that governs personal data of all types (General Data Protection Regulation (GDPR) – Official Legal Text, 2019) - especially voice data (Frost & Sullivan, 2022) - but the prioritization of data security and the vetting of the third parties that may, even with transparent consent, handle and process voice data. +● Inclusivity: working at all times to allow all to be heard – across languages, dialects, genders, ages, ethnicities, types and levels of disability. +● Accountability: maintenance of, and adherence to, highest ethical standards throughout the voice development, implementation, and operational value chain. +A next step for the Open Voice Interoperability Initiative (see Section Six, below) will be to translate the principles of these four papers into tangible, implementable technical guidelines and specifications. +The Work in Progress appendix to this paper (see section A.16, p. 69) contains an early proposal for an enforcement mechanism for privacy constraints. The approach consists of two mechanisms: (a) a mechanism to determine if data should be shared and with whom it is shared, and (b) a mechanism to monitor and control data after it is shared. +Page 43 of 57 + +SECTION FOUR: LESSONS FROM OTHER INTEROPERABILITY INITIATIVES +4.0 Introduction +As noted in Section One, the coming world of voice will be one of diversity, one marked by a multitude of voice assistants and agents. At present, however, the realm of general-purpose consumer-facing conversational assistants is one of singular, proprietary platforms that do not interact with one another. +As noted, the Open Voice Network envisions a future in which proprietary walls to content and connection are lowered, and users may seamlessly move from one assistant to another according to intent – a future in which voice operates broadly like the web, and not like apps on a mobile platform. +This section provides an overview of two important interoperability initiatives outside the Open Voice Network. We applaud –and are learning from – both. We also anticipate some level of adoption and adaptation of the concepts described here, as our work progresses. +4.1. Amazon Voice Interoperability Initiative (VII)™ +The Amazon Voice Interoperability Initiative (VII) enables the deployment of multiple conversational assistants on a singular endpoint device. Assessed with the Open Voice Network framework, it offers limited delegation – e.g., it does not allow the delegation of a conversation, but allows the user to leave one assistant and engage with another. +In a VII implementation, each assistant is activated via its own ‘wake word,’ enabling users to talk to the assistant of their choice in a secure manner by simply saying its name. For example, a user can interact with a specialized conversational assistant (such as Fridge) for specific refrigerator interactions, and leave the specialized conversational assistant and switch to a general-purpose conversational assistant (Amazon Alexa) for general purpose interactions. In +Page 44 of 57 + +this example, the Fridge conversational assistant would be designed with the vocabulary and skills specific to refrigerators and would bypass the vocabulary and skills that Alexa understands. Likewise, Alexa would not need to deal with the vocabulary and skills for the refrigerator. +The assistants interoperate with each other via the Multiagent experience (MAX) toolkit. The toolkit provides the MAX Library and a Sample Application that demonstrates the interoperability of Alexa and a second independent voice assistant. The MAX Library facilitates interoperability between voice assistants according to the guidance provided by the Voice Interoperability Initiative (VII) Multi-Agent Design Guide. + +Figure 5.1.1: Amazon Voice Interoperability Initiative +Within the VII guidelines, an assistant can transfer a user to a second assistant when it cannot directly fulfill a user request and is aware that the second assistant on the device can likely fulfill that request. No data or context is passed between assistants during a transfer, and the customer repeats their request directly to the second assistant without needing to say the wake word. +In addition, assistants interact with the device via Universal Device Commands. UDCs are commands and controls that a customer may use with any compatible assistant to control certain device functions, even if the assistant was not used to initiate the function. Examples +Page 45 of 57 + +include changing the volume of the device’s output speaker or stopping a sounding timer or music initiated from assistant B when assistant A is in control. +Comparison to the Open Voice Network proposal +● The focus of Amazon’s VII is on multi- assistant devices, i.e. several assistants accessed from one device. +● Delegation is limited to the assistants registered to the device. The Open Voice Network envisions a limitless number of potential voice-enabled destinations. +● Interoperability and assistant “discovery” is handled by the MAX toolkit and is limited to assistants registered and running on the device. In contrast, OVON allows for interactions between assistants potentially running exclusively on the cloud. +● In VII, assistants are discouraged to talk directly or pass information and context to each other. The user will need to repeat the queries to the second assistant. +● VII demands the presence of two running middlewares to mediate assistant interactions (the process running the MAX library and the device application and UDCs). +4.2. The Stanford Open Voice Assistant Laboratory (OVAL) Model +The Stanford University Open Voice Assistant Laboratory, under the direction of Professor of Computer Science Dr. Monica Lam, has proposed an open-source model for voice interoperability that enables mediation and partial delegation to multiple independent devices and services. (Lam et al., 2021). (Full disclosure: Dr. Lam is a valued advisor to the Open Voice Network, and the OVON is a financial supporter of the OVAL.) + +The Stanford OVAL model also follows the “Standardized Open Single Platform” model. The Stanford OVAL model gives developers the ability to collaboratively create “articles” through a device and services knowledge base known as “Thingpedia.” This empowers users to move from an initiating assistant to voice-enabled devices (i.e., a smart light bulb or smart factory sensor) and various services (a smart home system, a restaurant or retail website, or media properties such as Twitter or radio content). OVAL has demonstrated this model with several +Page 46 of 57 + +partners, using the OVAL “Genie” open-source conversational assistant as both mediator and initiator. +A strength of the Stanford OVAL model is that it allows the chaining and management of different articles combined in a single utterance (i.e., “When weather goes over 100F tweet out ‘oooh it’s hot’”). This also allows users to access multiple devices/services in a single request, although interoperability is restricted to the destinations, services, and use cases identified in Thingpedia. +Another significant strength of the OVAL model is that it enables users to enroll and save their devices/services. This allows users to interact with these devices/services securely and privately (without the need to authenticate separately and repeatedly). + +Page 47 of 57 + +Figure 5.2.1 Stanford OVAL architecture +The three core elements of the OVAL model: +ThingPedia +● ThingPedia is the knowledge base of what can be done by each device/service. +ThingSystem +● ThingSystem stores all device/service credentials for individual users allowing the platform to maintain privacy and security for the user. It runs the ThingTalk code that is provided by ThingPedia. +ThingTalk +● ThingTalk is the programming language that lies at the heart of OVAL platform. It gives the platform the power to connect IoT devices, web services, and database queries that are specified in ThingPedia. +Observations: +1. We perceive the OVAL model to be primarily one of single- assistant mediation. +2. If the OVAL model (with knowledge base Thingpedia and programming language ThingTalk) were widely adopted, it would deliver the capabilities of the Discovery and Location services identified above as a requisite for interoperability. +3. As the ThingPedia service expands, we fear that its ability to deliver discoverability may become strained. Currently Thingpedia has several hundred distinct “skills.” We must learn plans (UX, governance, etc.) for the scaling of the knowledge base to incorporate millions (or billions) of future skills. +Page 48 of 57 + +SECTION FIVE: FURTHER STUDY AND NEXT STEPS +This paper is but the first step on an important but challenging journey toward mediated and delegated voice assistance interoperability. +Visible next steps for the Open Voice Network Architecture Work Group include the following: +● The review, discussion, revision, and confirmation of the concepts shared in the Work in-Progress content in the Appendix to this white paper. +● The identification and assertion of messaging protocols that will allow independent voice assistants to mediate, delegate and serve as destinations for dialogs. With the encouragement of the Open Voice Network Steering Committee, the Architecture Work Group will first pursue adopting or adapting both existing standards and technologies that are now broadly used within the voice industry. +● Continued research as to mediated and delegated interoperability with existing and emerging voice and conversational AI services, including: +a) interactive Voice Response (IVR) systems inside corporate data security firewalls b) voice-enabled conversational bots +c) voice-enabled web applications +d) conversational AI implementations within enterprise software and processes e) enterprise-focused Conversational AI platforms. +● The development, with the OVON Voice Registry Work Group of the Technical Committee, of a standards-based approach to discoverability, findability, and location services. We anticipate publication of a separate document on this issue in the months ahead. +● The expansion of OVON messaging protocols to enable the sharing of multi-modal content. +Page 49 of 57 + +SECTION SIX: OPERATIVE VOCABULARY +Component: an identifiable part of a voice assistant or agent. A component provides a particular function or group of related functions. +Context: Information extracted from n prior utterances of the current conversation. This could include some or all of the following: information that has been inputted to outputted from, or inferred in Conversational Processors, and the information state of the Dialog Manager. Also known as ‘Conversational Context.’ +Conversation: a joint activity in which two or more assistants (human or automated) use linguistic forms and non-verbal signals (i.e., gestures) to communicate to achieve an outcome that meets a shared goal. +Conversation Event: a conversation event signals shifts in the conversation that may be acted upon. Such an event may occur at the beginning or ending of a Conversational Session, completion of a Conversation Processor, decoding of Conversation Information, changes to the state of a Conversation Endpoint, or changes to the status of a Conversation Stream. Any component with access to the system is allowed to generate a Conversation Event. +Conversation Facilitator: a component that coordinates communication between two or more Dialogue Systems and/or Processors during the course of one or more Sessions. This allows dialogue Systems and associated Processors to collaborate regardless of technology being used. Examples of Conversation Information include semantic, lexical, syntactic, and prosodic features. +Conversation Information Layer: an abstraction of a type of information in a Dialog System. A layer may be a specific type of acoustic, linguistic, non-linguistic, or paralinguistic features. Examples of layers would be Cepstral features, Phonemes, Intonation Boundaries, Words, Phrases, Turn Boundaries, Syllabic Stress, Discourse Move Type and specific Semantic representation schemes. +Conversation Processors: conversation information is encoded and/or decoded by one or more Conversational Processors. Conversational Processors may also take as input the output from another Conversation Processor. A Conversation Processor may generate Conversation Events and Conversation Streams. Example conversational processors include Automatic Speech Recognizers (ASR), Natural Language Processors (NLP), Dialog Managers, Text to Speech Synthesizers (TTS), etc. +Conversation Session: a particular conversation that consists of two or more Conversation Streams (see below) generated by two or assistants through one or more Conversational +Page 50 of 57 + +Endpoints. Sessions may be persistent, but they will often have a start-point and an end-point in time determined by one of the assistants or some other external event. +Conversation Stream: each Conversational Endpoint generates one or more Conversation Streams based upon the capabilities of the Endpoint and the preferences of assistant. A Conversation Stream is associated with a particular assistant and may include any media type including text, audio, video, and application UI events. +conversational assistant: a digital participant in a conversation. This may be an application with a consistent persona, such as Amazon Alexa, Google Assistant, the Target Google Assistant Action, a Facebook Messenger chatbot, or an IVR system at a bank or a human. conversational assistant is a common term utilized in Dialogue System research and university level instruction; the term often is used to describe a human participant in a conversation. For clarity of reference, however, the OVON will use the term "user" to identify a human participant. (See "user" below.) +Conversational AI: the set of technologies to enable automated communication between computers and humans. This communication can be speech and text. Conversational AI recognizes speech and text, understands intent, decipher various languages, and responds where it mimics human conversation. In some cases, it is also known as Natural Language Processing. +Conversational Context -See ‘Context,’ above. +Conversational Delegation: the passing of dialog layers and control between one Conversational Assistant and another to fulfill a user intent. The first assistant in the delegation sequence is the initiating assistant; the second is the destination assistant. +Conversational Endpoint: assistants conduct conversations using conversational endpoints; these may be a phone, mobile device, voice speaker, personal computer, kiosk, or any other device that enables an assistant to participate in a conversation. Endpoints may be referred to elsewhere as a "device" or a "channel." +Conversational Information Packets: information that relates to a specific period of time. Packets form the input and output of Conversation. +Conversational Mediation: the hosting of a dialog by a Conversational Assistant. In conversational mediation, the host assistant may fulfill a user intent by itself; it may access third-party data sources through API calls; or, it may introduce to the user a third-party application that is resident on the platform of the host assistant. A mediating assistant does not cede control, nor access to the data within the conversation. +Page 51 of 57 + +Conversational Platform: A group of technologies that are used as a base for one or more conversational assistants; also (see "Platform" below) a business model that harnesses and creates a large, scalable network of users and resources that can be accessed on demand. +Data: (per the Cambridge Dictionary): information, especially facts or numbers, collected to be examined and considered and used to help decision-making, or information in an electronic form that can be stored and used by a computer. Context (see above) is a subset type of the data accessed and used by the voice assistant system. +Dialog Manager (DM): handles the dynamic response of the conversation. It provides a more personalized response based on the action provided by the NLP to send back to the user. +Disambiguate: when the conversational platform hypothesizes two or more possible resolutions to a user utterance, it may ask the user for additional clarification or choose between the various interpretations to decide the user's correct intention. +Entity: a custom level data type and considered a concrete value to associate a word(s) in a query. It is a part of the structural machine translation. Also known as annotations. +Explicit Invocation: an invocation type where the user invokes the channel, and explicitly states a direct command to accomplish a specific task. The direct authority is to communicate directly to a registered voice application. +Implicit Invocation: an invocation type where the user invokes the channel by describing the target destination rather than naming the target. +Invocation: a part of the construct of the user's utterance during a conversation with a channel. An invocation describes a specific function that the guest wants, and solicits a particular response. +Intent: the identified action that the machine interprets based on the user's query. It is a part of the structured machine translation. Also known as a classifier. +Jobs to be Done: An approach to learning what will cause a customer to hire or bring your product or service into their life. +Natural Language Processing (NLP): a service and a branch of Artificial Intelligence that helps computers communicate with humans in their language and scales other language-related tasks. NLP helps structure highly complicated, unstructured human utterance and vice-versa. Natural Language Understanding is a subset of NLP that is responsible for understanding the meaning of the user's utterance and classifying it into proper intents. + + +Organization: a group of individuals brought together for a specific purpose, including the creation, transaction, and delivery of products or services. Examples would include a for-profit business, a not-for-profit group, or a government agency. +Platform: The collection of components (the environment) needed to execute a voice application. Examples of platforms include the Amazon and Google products that execute voice applications. +Query: user’s word requesting for specific function and expecting a particular response. +Speech-To-Text (STT): conversion of a representation of an utterance from audio to text.Also known as Automatic Speech Recognition (ASR). +Text-To-Speech (TTS): conversion of a representation of an utterance from text to audio. Also known as Speech Synthesis. +Technical Resource: it can be a publisher/developer. It can be a representative of an entity or independent party. Their role is to create an actual listing of the voice application. +Utterance: spoken or typed phrases. +User: a person who interacts with channels. +SECTION SEVEN: ABOUT THE OPEN VOICE NETWORK +The Open Voice Network (OVON) is a non-profit industry association dedicated to the development of standards for voice assistance transparency, consent, limited collection, and control of voice data that will make using voice technology worthy of user trust. In any reality, virtual or otherwise, we believe personal privacy should be respected as the default. The Open Voice Network operates as an open-source community within The Linux Foundation. It is independently funded and governed with participation from more than 120 voice practitioners and enterprise leaders from 12 countries. +The Open Voice Network community’s work is open source. We seek inclusive input and like to share our insights. At present, our work is focused in four areas: + + +● Interoperability, defined as the ability for conversational assistants to share dialogs (and accompanying context, control, and privacy) +● Destination registration and management, the ability of users to confidently find a destination of choice through specific requests, and for the providers of goods and services to register a verbal “brand” — similar to the Domain Name System (DNS) of the internet +● Privacy, with voice-specific guidance for both the protection of individual user data and that of commercial users +● Security, with a focus on voice-specific threats and harms. + +Please see our 2022 papers and support the Open Voice Network by visiting openvoicenetwork.org. +About The Linux Foundation +Founded in 2000, The Linux Foundation is supported by more than 1,000 members and is the world’s leading home for collaboration on open-source software, open standards, open data, and open hardware. Linux Foundation’s projects are critical to the world’s infrastructure including Linux, Kubernetes, Node.js, and more. The Linux Foundation’s methodology focuses on leveraging best practices and addressing the needs of contributors, users, and solution providers to create sustainable models for open collaboration. For more information, please visit us at linuxfoundation.org. +The Linux Foundation has registered trademarks and uses trademarks. For a list of trademarks of The Linux Foundation, please see its trademark usage page: www.linuxfoundation.org/trademark-usage. Linux is a registered trademark of Linus Torvalds. + + +Acknowledgements +This paper is authored by the Open Voice Network, with special thanks to the Architecture Work Group of the Technical Committee. Contributors: David Attwater, serving as Senior Research Scientist; Oita Coleman; Dr., Deborah Dahl, Senior Editor; Bruce Epstein, Co Moderator of the Architecture Work Group; Pradeep Gopal; Vineet Hingorani; Olga Howard; Carl Jahn; Kiran Kadekoppa; Dr. Jim Larson, Co-Moderator of the Architecture Work Group; Tobias Martens; Dr. Yaser Martinez-Palenzuela; Nick Myers; Shyamala Prayaga, Co Moderator of the Architecture Work Group; Elizabeth Robins; Dr. Dirk Schnelle-Walka; Nathan Southern; Jon Stine; Vadim Tarasevic; John Trammell; Boris Volfson. +We are grateful for the ongoing support of the Steering Committee of the Open Voice Network: Joel Crabb, Chair; Mirko Saul, Vice-Chair; Ali Dalloul, Bernhard Hochstätter, Doug Rogers, and Christian Wuttke. We wish also to thank and recognize individuals whose guidance, encouragement, and founding vision made this effort possible: Mike McNamara, Dan Cundiff, Kristi Dank, Maria Brinas-Dobrowski and Jay Kline; Ryan Steelberg and Sean King; Dr. Monica Lam and Jimmy Garcia-Meza; Reghu Ram Thanumalayan; Xuedong Huang; Bradley Metrock; Birgit Popp, Ulrike Stiefelhagen, and Anna Leschanowsky; Lawrence Lin. + + +SECTION EIGHT: REFERENCE LIST +Christensen, C., Hall, T., Dillon, K., & Duncan, D. (2016). Know Your Customers’ Jobs to be Done. Harvard Business Review, 94(9), 54–62. +https://hbr.org/2016/09/know-your-customers-jobs-to-be-done +European Data Protection Board. (2021, July 7). Guidelines 02/2021 on Virtual Voice Assistants. Retrieved August 11, 2022, from https://edpb.europa.eu/edpb_en +Fielding, R. T. (2022, June 1). RFC 9110: HTTP Semantics. RFC Editor. Retrieved August 11, 2022, from https://www.rfc-editor.org/rfc/rfc9110.html +Frost & Sullivan. Opportunities in the Conversational AI Market (2022, August); frost.com. Retrieved August 8, 2022 in advance of web publication. +General Data Protection Regulation (GDPR) – Official Legal Text. (2019, September 2). General Data Protection Regulation (GDPR). Retrieved November 8, 2022, from https://gdpr-info.eu +Haas, M. (2020, July 16). Understanding Conversational AI Tech. Interactions.Com. Retrieved August 11, 2022, from https://www.interactions.com/blog/technology/conversational-ai technology/ +IEEE Standards Information Network/IEEE Press. (2000). The Authoritative Dictionary of IEEE Standards Terms (IEEE 100), Seventh Edition (7th ed.). Institute of Electrical and Electronics Engineers (IEEE). +Johnston, M., Baggia, P., Burnett, D., Carter, J., Dahl, D., McCobb, G., & Raggett, D. (2009, February 10). EMMA: Extensible MultiModal Annotation markup language. World Wide Web Consortium. Retrieved August 11, 2022, from https://www.w3.org/TR/emma/ +Lam, M., Landay, J., & Manning, C. (2021). Launching a World Wide Voice Web. Stanford University. + +McGlashan, S., Burnett, D., Carter, J., Danielsen, P., Ferrans, J., Hunt, A., Lucas, B., Porter, B., Rehor, K., & Tryphonas, S. (2004, March 16). Voice Extensible Markup Language (VoiceXML) Version 2.0. World Wide Web Consortium. Retrieved August 11, 2022, from http://www.w3.org/TR/2004/REC-voicexml20-20040316/ +∞ +2022.08.18/08:50 + From 8bfa06a460be3fd530658399a1837e0d5f671c25 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:01:40 -0500 Subject: [PATCH 27/35] Update and rename Interoperability of Conversational Assistants to Interoperable Dialog Event Object Specification 1.0 --- ...eroperability of Conversational Assistants | 698 ------------------ ...able Dialog Event Object Specification 1.0 | 516 +++++++++++++ 2 files changed, 516 insertions(+), 698 deletions(-) delete mode 100644 specs/Interoperability of Conversational Assistants create mode 100644 specs/Interoperable Dialog Event Object Specification 1.0 diff --git a/specs/Interoperability of Conversational Assistants b/specs/Interoperability of Conversational Assistants deleted file mode 100644 index 516cd7d..0000000 --- a/specs/Interoperability of Conversational Assistants +++ /dev/null @@ -1,698 +0,0 @@ - -2022.08.12 -Version 1.0-B -Interoperability of Conversational Assistants A NEW APPROACH -The Open Voice Network -Architecture Work Group of the Technical Committee -12 August 2022 - - -Executive Summary -The Open Voice Network (OVON) is an open-source community of the Linux Foundation, dedicated to developing the technical standards and usage guidelines for the emerging world of conversational Artificial Intelligence and, within that, voice assistance. -The future of conversational Artificial Intelligence is a future of diversity – of infrastructure providers and platforms, of enabling technologies, devices, and enterprise use cases. Conversational AI – and within conversational AI, voice assistance -- will be a primary interface to the internet, personal transportation, smart environments (domestic and enterprise), and the immersive digital world of gaming and augmented and virtual experiences. -To reap the greatest economic and social value of conversational assistance, we believe that voice must work like the web, enabling users to access any voice-enabled content destination regardless of platform. Interoperability – the sharing of dialogs between conversational assistants and agents of different technological parentage – is an essential capability to reach the value of what will become a Worldwide Voice Web. It must also be worthy of individual and enterprise user trust; readers will note references here to OVON work in data privacy, data security, and ethical use. -This paper is a step on the path to conversational assistance interoperability. In this paper, we assert that the proper model for conversational assistance interoperability is a human one – and suggest that humans generally resolve questions with one another through a process of either mediation or delegation. Using this framework, we identify the jobs to be done for priority constituencies, the architectural patterns that must be resolved, lessons learned from other work, and a path forward to an ecosystem of standards-based OVONICA – Open Voice Network Interoperable Conversational Agents. - -ABSTRACT -This is a publication of the Open Voice Network (OVON, www.openvoicenetwork.org), a non profit industry association operating as an open-source community of the Linux Foundation. It asserts that, to realize the full economic and societal potential of conversational assistance, conversational assistants must not only mediate human-to- assistant conversations – e.g., host the conversation and obtain relevant information to fulfill user intents – but also delegate conversations to other assistants, and in doing so, pass textual, acoustic and contextual data, as well as privacy and security controls. -To meet this vision, the Open Voice Network proposes an approach to interoperability between conversational assistants – specifically, the sharing of multi-layered dialogs between assistants and assistants of differing infrastructures through the standardization of communication protocols by which autonomous assistants and assistants collaborate to achieve a common goal. -This paper also suggests areas for further research, as well as next steps for the proposal and testing of communication protocols that may be built from existing technologies and universally accepted standards. -At the core of this work is this firm belief: we are in the early days of conversational assistance, and the future will be one of stunning diversity -- a multiplicity of voice-enabled user end points, content and infrastructure providers, voice-enabled destinations (measured, perhaps, in the billions), industry ecosystems such as transportation and smart homes, and organizational/enterprise value propositions. -This paper is the creation of the Architecture Work Group of the Technical Committee of the Open Voice Network. The Open Voice Network is a trademark of The Linux Foundation. Other trademarks referenced in this report are the property of their respective owners. - - -TABLE OF CONTENTS -1.0 Introduction 6 1.1 Intended Readership 7 1.2 Purpose of this White Paper 7 1.3 The Open Voice Network and Interoperability of Conversational Assistance: Why and What 8 1.4 Going Forward: Your Participation 10 1.5 A User’s Vision: An Interoperable World for Today and Tomorrow 12 1.6 Architectural Aspirations 14 1.7 Boundaries 15 1.8 Interoperability Defined for Conversational Assistance 18 1.9 Core Requirements for Interoperability of Conversational Assistants 19 1.10 Interoperability Intentions of the Open Voice Network 22 -SECTION TWO: A FOUNDATION FOR ANALYSIS -- THE TECHNOLOGY OF CONVERSATIONAL ASSISTANCE 23 -2.0 Introduction 23 2.1 Conversational AI and Voice 23 2.2 Conversational Assistants and Platforms 24 -2.2.1 Conversational Assistants and Agents 25 2.3 Platforms and Content: Existing Standards 26 2.4 Conversational Architecture 26 2.5 The Diversity of Today’s Conversational Assistance 29 -SECTION THREE: AN APPROACH TO VOICE ASSISTANCE INTEROPERABILITY 30 3.0 Introduction 30 3.1 Modeling Interoperability on How People Communicate 30 3.2 Interoperability: Mediation and Delegation 31 3.3 From the Human Model to Standards: OVON Interoperability Objectives 33 - - -3.4 Architectural Patterns Under Investigation 34 3.4.1 Dialog Delegation and Give-Back 336 3.4.2 Delegation Request and Conversation Context 36 3.4.3 Dialog Interaction PayLoad 37 3.4.4 Dialog Component Interfaces 37 -3.4.4.1 An Illustrated Example: Pat Goes Shopping 39 3.4.5 Discovery and Location 40 3.4.6 Sharing and Protection of Data 41 -SECTION FOUR: LESSONS FROM OTHER INTEROPERABILITY INITIATIVES 43 4.0 Introduction 43 4.1 Amazon Voice Interoperability Initiative (VII) 44 4.2 The Stanford Open Voice Assistant Laboratory (OVAL) Model 45 -SECTION FIVE: FURTHER STUDY AND NEXT STEPS 48 SECTION SIX: OPERATIVE VOCABULARY 49 -SECTION SEVEN: ABOUT THE OPEN VOICE NETWORK 53 About The Linux Foundation 54 Acknowledgements 54 -SECTION EIGHT : REFERENCE LIST 556 WORK-IN-PROGRESS APPENDIX 56 -Page 5 of 57 - -PREFACE -The Open Voice Network (OVON) was founded by users of conversational assistance. Its purpose is to develop and drive adoption of technical standards and usage guidelines that will make the emerging world of conversational assistance worthy of user trust. -This document represents the work of but one of several OVON technical work groups – all under the aegis of the Open Voice Network Technical Committee. Separate, yet related initiatives are at present exploring issues of data protection and data security for conversational assistance, assistant-agent destination and location services, voice-centric authentication, and synthetic voice. Other envisioned work streams await resources. -To date, the work has been developed on parallel tracks. As we go forward into Q4 2022, the streams will begin to merge – informing and being informed by each other. -We look forward to collaborating with you. -SECTION ONE: RATIONALE, SCOPE, AND DEFINITION -1.0 Introduction -This section speaks to the Open Voice Network’s reasons for pursuing interoperability of conversational assistants, the boundaries of our work, and a 1.0 definition of conversational assistant interoperability. -Page 6 of 57 - -1.1 Intended Readership -This paper was published for public evaluation and criticism on 12 August 2022. It seeks to address three global audiences, each with a stake in the future of voice: -● Enterprise content creators and communicators (business-to-consumer and business to-business) -● Stakeholders and decision makers within the voice technology industry, including potential architectural partners -● Current and prospective participants and sponsors of the Open Voice Network. -1.2. Purpose of this White Paper -This White Paper seeks to: -● Identify the architectural elements required to facilitate interoperability of conversational assistants -● Prioritize these elements for standardization by the OVON Architecture Work Group ● Foster discussion with the constituents noted above in section 1.1. -In this White Paper, we seek to remain neutral with respect to competing architectural directions, and to reveal the underlying concepts/issues that exist regardless of architectural choice. -This document is intended to be read alongside the Open Voice Network Technical Master Plan. In the short term these two documents may contradict one another as ideas are developed and tested. Such conflicts will be resolved as they emerge. -This paper replaces a previous document entitled ‘Architecture Design.’ -1.3 The Open Voice Network and Interoperability of Conversational Assistance: Why and What -Page 7 of 57 - -The Open Voice Network (www.openvoicenetwork.org) is an open-source community of the Linux Foundation dedicated to the development of the standards and usage guidelines for the emerging world of voice assistance. -It enjoys the regular participation of more than 200 volunteers from 13 nations and 5 continents. The OVON community includes participants from leading voice platforms and infrastructure providers; large enterprises that use voice technologies in customer service and operations (from industries such as retail, healthcare, telecommunications, and financial services); marketing and consulting firms; and more than 40 independent voice development and services companies. -As a technology-neutral, nonprofit organization, the Open Voice Network occupies a unique and strategic position within the voice technology industry. Our sponsors and participants witness a growing diversity of underlying voice technology (speech recognition, language understanding, dialog modeling), conversational design paradigms, labels to represent semantic content, and endpoints (from smart speakers to voice-enabled web pages and immersive games). They also see the rapid growth and evolution of voice value propositions, especially for organizations and across vertical industries. -This diversity points to a significant total available market (TAM) for conversational AI as both an interface and a source for data-fueled insights. Looming before us – as first envisioned by Dr. Monica Lam and colleagues at Stanford University – is an interoperable Worldwide Voice Web (WWVW) (Lam et al., 2021)., where the spoken word offers an open, standards-based interface to billions of voice-enabled media, enterprise, website, transportation, smart IOT environment, and metaverse destinations. -Given the economic and societal value that will be created by an interoperable Worldwide Voice Web, the Open Voice Network commissioned the Architecture Work Group of its Technical Committee to research and recommend architectural options for open, standards based interoperability for conversational assistance. -Using a technologically and architecturally neutral eye, the Architecture Work Group reviewed existing technology standards and existing and proposed voice architectures and messaging protocols. The Work Group ascribed to the Harvard Business Review’s “Jobs to Be Done” methodology (Christensen et al., 2016)., and examined in detail the current and future needs of four key voice constituents: -Page 8 of 57 - -● Consumers of voice-enabled experiences -● Enterprise content providers -● Technical innovators of dialog systems -● Voice system infrastructure providers. -Through these efforts, the work group identified the opportunities and challenges of conversational assistant interoperability, and a series of foundational concepts. -We are deeply grateful for the contributions of the many Open Voice Network participants and sponsors referenced on page 55. These individuals gave freely of their time and intellect to develop this paper and the plans to take it forward. -The Open Voice Network: -● Believes that conversational assistance will realize its economic and societal value when it is interoperable, like telephony or the worldwide web (WWW). -● Sees in today’s rapidly growing enterprise investment (Frost & Sullivan, 2022) in conversational assistance the seeds of a rich ecosystem of independent, purpose- and brand-specific voice assistants, an ecosystem that will bring desired content and experiences to the users of general-purpose, proprietary conversational assistants. -● Recommends the development and industry-wide adoption of standardized communication protocols between conversational assistants operating on different platforms. -● Asserts that conversational assistants must both mediate dialogs (e.g., act as a host, and acquire desired information to fulfill the user intent) and delegate dialogs (e.g., act as an initiator of communication, and bridge to other assistants so that the user intent can be fulfilled.) In addition, conversational assistants must also be identified and authenticated to serve as a destination of a delegating assistant or agent. -● Envisions a model of conversational assistance interoperability in which autonomous assistants and agents collaborate to achieve a goal together. In so doing, they will use standardized ways to share information with each other and operate at various, negotiated levels of trust and information sharing. This will lead to the development, test, and proposal of standards for: -Page 9 of 57 - -o The way in which a spoken assistant name can be used to find an associated assistant -o The way in which basic linguistic information is shared between dialog assistants o The way in which immediate linguistic context and history is shared between dialog assistants -o The way in which control is handed between dialog assistants -o The way in which dialog assistants negotiate trust. -In addition, the Open Voice Network -● Will neither develop nor recommend the standardization of platform components or the format of content used to figure these components. -o We do not believe that a standard methodology for expressing and describing conversational interaction is possible or desirable. Attempts to describe how conversations can or should be modeled will quickly become outdated as conversational interaction innovates and evolves. -In the coming months, the Open Voice Network will develop, test, demonstrate, and propose to existing standards bodies a set of standardized communication protocols that will enable assistant-to- assistant voice interoperability. -1.4 Going Forward: Your Participation -In keeping with the practices of the Linux Foundation, the Open Voice Network works in an open, communal manner. We seek contributions from every corner and region of the conversational AI ecosystem. -Readers of this paper are invited to -● Propose corrections and additions to this paper through the Open Voice Network GitHub repository (https://github.com/open-voice-network) or in the White Papers section of the Open Voice Network website. -(https://openvoicenetwork.org/white_papers). We are especially interested in criticism and comment regarding -o Concepts that are not clearly defined -Page 10 of 57 - -o Requirements that you feel are missing -o Requirements that you feel are unnecessary -● Contribute usage scenarios (similar to those in the paper) that challenge the concepts within this paper, and perhaps point us to new or revised requirements for the interoperability of conversational assistance. -You are also welcome to join the weekly interoperability development meetings of Architecture Work Group of the Open Voice Network’s Technical Committee. One-hour sessions are scheduled Tuesdays at 17:00 CET, 11:00 Eastern, and 08:00 Pacific. Conferencing details are found at https://openvoicenetwork.org/calendar. -Page 11 of 57 - -1.5 A User’s Vision: An Interoperable World for Today and Tomorrow -Today: Pat Shops by Voice as Humans Do -Pat has finished the day’s labors and is driving home. It’s late. A request to Pat’s automobile voice agent starts a music playlist that brings back memories of school days – which, with a start, reminds Pat that a long-time friend from school, now in town for a conference, will be coming to dinner tomorrow. -What to prepare? Perhaps a favorite seafood stew. Crab, mussels, shrimp, snapper in a tomato broth with wine? Pat asks the automobile-based voice agent to connect to the grocer that knows Pat’s preferences – and, in an instant, Pat is directly connected, and speaking to the grocer’s conversational assistant. -A recipe is identified. A shopping list is created -- items, quantities – based upon availability of overnight shipments from seafood suppliers. Wines are recommended (a Zinfandel and Sauvignon Blanc) and chosen. -Pat sighs with relief. And then – right before completed list is transferred to Pat’s mobile grocer app for authentication, pick-up scheduling, and payment -- Pat remembers that the school friend has a chronic medical condition (in this case, diabetes.) Pat quickly asks the grocery assistant if the seafood stew will be healthy or unhealthy. The assistant, in turns, informs Pat that it is directly connecting to a conversational capability developed specifically to recommend food choices for individuals with diabetes. -From the conversational capability, Pat learns quickly that the seafood stew will be a healthy choice for dinner with the school friend. A quick request to the conversational capability, and Pat is sent back to the grocery assistant – and the shopping list is confirmed. -Page 12 of 57 - -Tomorrow: Pat Plans Travel by Voice as Humans Do -Pat plans to travel to an international conference and stay for the weekend that follows. Pat needs a visa, airline reservations, and information about attractions for anticipated free time. Pat’s personal conversational agent is named “Butler,” and was developed by a third party using a mix of proprietary and open technologies. -Pat begins trip planning by pressing a button on a smartphone to begin conversation with Butler. Butler is local to Pat’s phone and has permission to access Pat’s personal data. Butler invokes a voice passport authentication capability to verify that the speaker is Pat and uses Pat’s smartphone for second-factor authentication. -Pat informs Butler of the conference in a major city on another continent. Pat has been asked to deliver a keynote address. A visa may be required to attend. Butler uses a voice destination and location service (similar in purpose to a Domain Name System) to discover the conversational assistant of the passport-visa office of the destination country. -The visa conversational assistant provides Butler with the latest visa requirements and guidance on how (and when) to apply; Butler forwards the guidance and application URL to Pat’s digital account and schedules a reminder to ensure its submission. At the same time, Butler has explored flight options by searching among highly rated travel services. -Butler selects AirWithFlair, a specialty conversational agent and capability for artists and musicians., and mediates a connection to obtain a list of travel options (and prices) for Pat. Butler - knowing that Pat is a loyal customer of the Royalty Hotel chain - delegates the dialog to the Royalty conversational assistant, which recognizes Pat and immediately identifies a location, room type, and pillow choice for the potential conference visit. -Page 13 of 57 - -1.6. Architectural Aspirations -The following architectural aspirations guide our thinking and our actions: -1. We recognize that the material and financial resources required to create, operate, and maintain conversational AI and voice assistant systems are significant. Our architectural direction will allow for sharing the cost of those resources, but we will not ourselves specify how that needs to occur. -2. There are several competing architectures for control interfaces between conversational AI platforms, assistants, and agents. OVON will evaluate architectures, produce recommendations, and attempt to remain neutral with respect to the choice until and unless making a choice becomes required for completing our specifications. -3. Each of the dominant proprietary platforms within the existing voice market has an installed base of platform-specific applications. At the same time an ecosystem of third party voice content creators has emerged outside the proprietary platforms. To maximize the value of our work and speed global adoption, we will seek to establish a platform-neutral approach to interoperability – while adopting (as noted above) common concepts and ideas. -4. The lifespan of a standard is much longer than the lifespan of the physical components subject to the standard. The specifications and guidance which we will produce may be informed by existing design patterns but must not be limited to those patterns. -5. In the immediate, we aspire to: -a. Outline the different architectural approaches that could support -interoperability. -b. Outline the advantages and disadvantages for each with respect to the OVON constituents and the Jobs to Be Done. -c. Recognize that several viable architectural approaches to interoperability may coexist and that OVON needs to remain responsive to the directions taken by the market. -d. Identify platform-neutral standardization that will enable both the new emerging third-party ecosystem and the applications resident on proprietary voice platforms. -Page 14 of 57 - -When you’re exploring the internet, do you shut your browser down to open a new web page? -Why, then, would you want to do that with voice? -Royal O’Brien, Linux Foundation, July 2022 -1.7 Boundaries -The Open Voice Network defines the boundaries of this work using the “Jobs to Be Done” approach (Christensen et al., 2016)., singling out four groups of individuals for whom interoperability will create incremental and sustainable value. Identification and clarification of the needs and desired experiences of each group allows the OVON to define the work required from the Open Voice Network to fulfill those needs and desires. -Though Open Voice Network-delivered standards will benefit the voice industry at large, our work suggests that four distinct groups of individuals are central to the question of interoperability: (1) consumers of enterprise content who wish to use voice enabled technologies, (2) enterprises who provide content and wish to provide voice enabled experiences, (3) technical innovators of dialog systems who wish to create the systems for the enterprises to use, and (4) voice systems infrastructure providers who wish to create and pursue commercial opportunities within a free and open environment. -Page 15 of 57 - - -Figure 1.7 Benefactors of conversational assistant interoperability; the value ecosystem -Together, these groups constitute a tiered ecosystem that will flourish in a world of conversational assistant interoperability. -Consumers of Voice Enabled Experiences are at the top of the tier. These are the end consumers of voice enabled experiences. OVON-developed standards will enable them to engage in voice experiences in web environments that currently do not provide voice experiences. The “Jobs to be Done” that we are supporting for this group include the following tasks: -● Help me continue my voice task as I switch voice devices (home, phone, auto) ● Help me avoid restricted access to data in all internet-connected, voice-enabled systems ● Help me opt-out of data protection when it makes sense. -● Help me avoid restrictions on vendor choice. -Enterprise Content Providers rank next. These are the businesses and organizations that offer applications on proprietary, general-purpose consumer platforms, or do not deliver voice experiences to their customers outside the organizational security firewall. Together, enterprise content providers represent a significant third-party content and voice-centric services ecosystem. “Jobs to be Done” for enterprise content providers are: -● Help me avoid building separate applications and services in order to use all global voice platforms and web / IoT / metaverse-based voice-based destinations -● Help me grant all potential constituents free and unfettered access regardless of the -Page 16 of 57 - -constituent's "home" voice platform -● Help me ensure that all my potential constituents can freely and directly connect with my voice applications and services, regardless of home platform -● Help me enable the privacy of all dialog data (text, acoustic, semantic) shared between my applications and services and my constituent when explicitly requested to do so (however that may be) -Technical Innovators of Dialogue Systems are the technology-centric and consulting organizations that will create customer service, web, or operational environments for enterprise content providers. These innovators develop the different technology solutions that work together to enable voice experiences. They will build the pieces of the modularized future we envision. “Jobs to be Done” for technical innovators include: -● Help me avoid creating new voice-based core business processes for transactions ● Help me avoid creating new/redundant messaging protocols for sharing of dialog, data, and controls between platforms, agents, and assistants -● Help me give customers the choice to protect their voice data -● Help me secure my customer's voice data -● Help me avoid creating new specifications /processes for voice authentication, attestation, and authorization -● Help me get to content and consumers without dependence upon Big Tech proprietary platforms -Of note: while there are “Jobs to Done” for each of the constituent groups – and each plays an important role in an interoperable ecosystem – all will rely on the work of the Technical Innovators. These individuals will use Open Voice Network standards to do their work; the OVON expects to provide significant support to assist them in accomplishing their jobs to be done, each of which require a specific experience. These are the experiences we will create for the technical innovators. -Technical Innovator Jobs to Be Done Corresponding Technical Innovator Experiences -Help me continue my voice task as I switch voice devices (home, phone, auto) Implement OVON standard protocols and specifications to enable business processes for interaction -Help me avoid creating new/redundant messaging protocols for sharing of dialog, data, and controls between platforms, agents, and Select and plug and play messaging protocols - - -Page 17 of 57 - -assistants -Help me give customers the choice to protect their voice data My customer elects to protect their voice data -Help me secure my customer's voice data I match security level per customer's level of data protection -Help me avoid creating new specifications /processes for voice authentication, attestation, and authorization I learn OVON Standards so that I know how it should work -Help me get to content and consumers without relying on the proprietary Big Tech I know where to go and what to do to manage proprietary Big Tech platforms - - -Voice Systems Infrastructure Providers provide technologies used by content creators and technical innovators to deliver voice-based content and services to consumers of voice-based experiences. In an open, standards-based, interoperable world, they will enable a stable voice ecosystem. Technical innovators will rely on these open infrastructure offerings to create voice technology solutions that will be used by enterprise content providers. “Jobs to be Done” that we are supporting for voice system infrastructure providers are the following: -● Help me create an accessible environment -● Help me create an environment where seamless interaction is commonplace ● Help me create an environment where scalability is empowered -Two additional jobs to be done were identified during working group consultation with a major infrastructure provider: -● Help me maintain ownership of data -● Make sure that when my infrastructure is used (work is being done by me), that my work gets recognized. -1.8. Interoperability Defined for Conversational Assistance -At the most basic level, interoperability means “the ability of two or more systems or components to exchange information and to use the information that has been exchanged.” IEEE Standards Information Network/IEEE Press (2000). The OVON adapts this definition to the conversational assistance ecosystem as follows: -Page 18 of 57 - -Interoperability is 1) the ability of conversational assistants of diverse parentage to collaborate seamlessly to achieve a goal together, using standard ways to share information with each other, operating at various levels of trust and information sharing, and operating independently of the physical devices being used; and 2) the ability of conversational AI-driven content, assistants, and service providers to reach target audiences across the diversity of platforms, general purpose assistants, and physical access devices (endpoints). -Conversational assistants can be fully or partially interoperable. -1.9 Core Requirements for Interoperability of Conversational Assistants Conversational assistants must: -● Be enabled to fulfill user intents through mediated and delegated communication. ○ In mediated assistance, the user interacts with a single conversational assistant. The mediating conversational assistant hosts the conversation; it never cedes control of the conversation. To fulfill user intent, the mediating assistant may obtain information from third party sources or introduce an application resident on its platform. -○ In delegated assistance, the user interacts with one or more independent conversational assistants using standardized messaging protocols. The hosting conversational assistant initiates the conversation; it passes control to a delegate assistant (or assistants). -● Be enabled to serve three roles to fulfill user intents: as 1) a host of applications that mediates content to fulfill user intent, 2) an initiator of a connected dialogue that delegates fulfillment of intent to one or more independent assistants, or 3) a destination that is delegated to, and as such, assists the initiator in fulfilling user intent. -Page 19 of 57 - - Figure 1.9: Core Requirements for Conversational Assistants -As an initiator and as a destination, an assistant must offer six types of interoperability, as shown in the grid below. -Type of -Interoperability Use Example Benefit -Share dialog -fragments -Dialog fragment = portion of a dialog pattern, a piece of interactive -fragment, a chunk of interaction -specification Conversational -assistant uses -fragments from -one or more other conversational -assistants to -achieve a dialog goal – user thinks is talking with one assistant A delivery service conversational assistant (with Pat’s consent) uses the location collection dialog -fragment from a payment -conversational assistant to collect Pat’s address information. The fragment may also be identified so it can be reused. Application -developers save time by reusing -dialog fragments from other -conversational -assistants. - - -Page 20 of 57 - -Share personal data -Personal data = information which the user wants to maintain across several sessions. -A conversational assistant specifies how data is shared with another -conversational assistant. -The value in the Amount Due slot in the Shop conversational assistant is copied into the Payment Amount slot in the Payperson conversational assistant. -User controls how personal data can be shared with other -conversational assistants, saving users from -reentering the data value. - -Share -conversational context -Conversational -context = -information which is shared between a user and a -conversational -assistant over a (tbd) period of time. With the user’s permission, -conversational -context may be shared with other conversational -assistants. Relevant context (of the user, of the dialog) established on one assistant is shared with -another -conversational -assistant. - The context (history) of actions in the Shop voice assistant is shared with the Payperson voice assistant. The Payperson conversational assistant is aware of information collected and generated by the Shop voice assistant. Enables -conversational -continuity -conversational -continuity = the -principle of making sure that all details in one voice -assistant are -consistent with the details in another voice assistant -Transfer control -Transfer control = determine which conversational -assistant is invoked and becomes -active. -Determination of which assistant has “the floor,” which assistant has the Transfer control from one -conversational -assistant to -another While shopping for groceries -through a conversational assistant, Pat switches to the Payperson voice assistant to pay for the groceries. Matches the -natural thought -processes of -human users. - - -Page 21 of 57 - -permission to -speak. -Share -components (This is an aspirational goal) -Components = -components of a conversational -assistant, -including but not limited to ASR, -TTS, NLP, NLG. The developer -specifies which -components are used within a -dialog. A voice developer may desire to use sophisticated natural language processing rather than a simple natural language processor provided by the conversational assistant being used. -Other examples of component sharing include (a) using different ASRs for speakers of American English and British English, (b) performing natural language -translation between national languages such as English/German or English/Mandarin, (c) providing access to algorithms that perform explicit query or implicit query (a.k.a. voice search). Developers can -develop custom assistants through the use of -standards-based, best-of-breed -technologies. -Share endpoints -Endpoint = -hardware or -software through which the user -speaks and listens to conversational assistants User controls -where input may come from or -output may be -directed to. The user may begin a request from a device with a public -microphone/speaker and later request to continue with a private device. -The user may ask for output to be directed to a different -conversational assistant or a -different device, such as a printer, or may ask for input to come from some shared data. User can access -any conversational assistant using any endpoint - - -Page 22 of 57 - -1.10 Interoperability Intentions of the Open Voice Network -The Open Voice Network seeks to enable mediated and delegated communication by conversational assistants within an ecosystem of general purpose, consumer-centric conversational assistants, and independent, purpose- and brand-specific conversational assistants and assistants. To make this happen, OVON will develop messaging protocols and a destination management and services system, to standardize the way in which: -● a spoken assistant name can be used to find an associated assistant -● basic linguistic information is shared between dialog assistants -● immediate linguistic context and history is shared between dialog assistants ● control is handed between dialog assistants -● dialog assistants negotiate trust. -SECTION TWO: A FOUNDATION FOR ANALYSIS -- THE TECHNOLOGY OF CONVERSATIONAL ASSISTANCE -2.0 Introduction -The Open Voice Network grounds its research into conversational assistant interoperability in current operative definitions of conversational Artificial Intelligence, as well as the normative components and architecture of contemporary voice assistance. In this section, intended for -both subject matter expert and lay audiences, we share our understanding of those components and architectures. -2.1 Conversational AI and Voice -Voice assistance is but one form of conversational Artificial Intelligence, or conversational AI. -Page 23 of 57 - -Citing Interaction.com (2022): “Conversational AI is the set of technologies behind automated messaging and speech-enabled applications that offer human-like interactions between computers and humans. -“Conversational AI can communicate like a human by recognizing speech and text, understanding intent, deciphering different languages, and responding in a way that mimics human conversation.” (Haas, 2020). -Conversational AI solutions can be offered through text and voice modalities. Voice solutions – referred here broadly as voice assistance – can be delivered as an audio-only interface, or one in which voice assistance works in a complementary way with screen-based information. The latter is often described as multi-modal voice assistance. -This paper speaks primarily to the sharing of voice-based elements within a dialog. However, the Open Voice Network recognizes clearly the importance of multi-modal voice assistance (especially in an enterprise context) and will extend future editions of this work to multi-modal interoperability. -2.2. Conversational Assistants and Platforms -The Open Voice Network draws a clear distinction between conversational assistants and conversational platforms. -Conversational assistants function as conversational user interfaces. They engage in conversation and are identifiable by users as conversational actors. A conversational assistant typically has an assigned identity, an identifiable voice, and/or a name: for instance, Alexa, Siri, or Magenta assistants. The user recognizes a given conversational assistant as a singular entity and refers to that assistant with a pronoun (he, she, it, they) or by an assigned name. The conversational assistant may also refer to itself using a name or a pronoun. -Conversational assistants often respond to queries in the form of web searches. The response typically describes how the assistant locates the delivered content. For instance, one might ask, “Assistant, what is the weather like in Brooklyn, NY today?” To which the assistant responds, “Okay, here is something I found on weather.com . . . “ -A conversational platform is the software to which the assistant connects, and that interprets the input from the user. The assistant responds through an endpoint, which is typically a piece of hardware, such as a smart home display, a television remote control, an automobile smart dashboard, a smartphone, or a smart speaker. -Page 24 of 57 - -There are endpoints and platforms currently associated with proprietary interfaces, such as Amazon Alexa or Google Assistant. They handle wake words (“Hey, Google…!”), audio coding and streaming, and graphical events and content. -If a given platform provides generic access to content, it may have a host voice assistant, and give the user the option to ask for a delegate assistant. In other words, the host can provide a gateway to delegate content. -With the exception of assistants written in the VoiceXML standard (McGlashan et al., 2004)., delegate assistants are typically authored for a given platform by third party companies. -Examples of End Point-Platform-Assistant Pairings -End Point Platform Host Assistant -Echo Dot Amazon Alexa -Google Home Google (Hey) Google -Magenta Mini Deutsche Telekom Magenta - - -Names and brands may be the property of others. -2.2.1 Conversational Assistants and Agents -For the purposes of this paper, the terms “assistant” and “agent” both refer (as noted above) to conversational AI user interfaces that are perceived to be a single conversational actor, operate on behalf of the user, and operate in accord with a conversational platform. Both assistants and agents have an addressable name, continuity of knowledge, and a bounded identity. -Both assistants and agents have the ability to fulfill user intents, mediate and delegate dialogs, and serve as a destination. Agents will also operate independently on behalf of the user. In the worldwide voice web, there will be specific-purpose conversational assistants and agents, and general-purpose conversational agents. specific-purpose conversational assistants and agents, and general-purpose conversational agents (McGlashan et al., 2004). -Page 25 of 57 - -2.3. Platforms and Content: Existing Standards -Most conversational systems separate the platform from the content that is executed on that platform. In Web Models - a common structure - the content itself is typically either static or active web content. The platform is roughly analogous to a browser; the content is stored on web servers and accessed by standard web protocols (e.g. HTTP) (Fielding, 2022). -A set of related standards proposed by the W3C Voice Browser group has successfully created standardized formats for describing content for spoken dialog, speech recognition grammar, semantics, text-to-speech markup and pronunciation (McGlashan et al., 2004). This content can be interpreted by a W3C compliant voice browser. -The W3C voice-browser group standards have served the telephony community well but have not yet been adopted by the leading general-purpose conversational platforms, nor by the many creators of text-centric chatbot platforms. One notable exception is speech-synthesis markup, which has been widely adopted (McGlashan et al., 2004). -This new wave of voice assistant platforms also uses a web-content model to support third party dialog content. The formats for such content are mostly published but remain proprietary. Examples include Amazon Skill Definitions and Google Actions. Content for this new wave of platforms utilizes standard web protocols such as HTTP or OAuth (Fielding, 2022), but also strongly promotes or mandates proprietary infrastructure and authoring tools. -2.4 Conversational Architecture -Below is a schematic view of a generic conversational assistant system. -The Open Voice Network bases its explorations of voice assistance interoperability upon this understanding of system components and architecture. -Page 26 of 57 - - Figure 2.4 Typical Components of a Conversational Assistant -Figure 2.4 illustrates the schema of traditional conversational architecture, along with the various components that are collectively enlisted to enable two-way communication with the user. -The architecture typically operates as follows, as per the above diagram, from left to right: -The user begins (far left) with an endpoint, which can be located on any piece of digital hardware. Today’s endpoint options include conversational assistant implementations on devices ranging from smart speakers, smartphones, desktop or laptop computers, and smart home systems, to automobile dashboards, television remote controls, and kitchen appliances – the list is endless. Endpoints provide a software or hardware interface between the user(s) and the Conversational Platform, which includes all the components within the encircled figure above. An example of a simple endpoint is an old-fashioned telephone (often referred to as the -Page 27 of 57 - -plain old telephone system, or POTS). A sophisticated endpoint might be an immersive VR environment.1 -Endpoints engage with the user via audio, text and visual media. Platforms interpret content that embodies voice assistants and enables them to interact with users via the endpoint. -The diagram above depicts three conversational components for exchanging information between the endpoint and the interaction manager. Let’s explore these as listed from top to bottom in Figure 2.4. -The most traditional form of user-assistant communication involves Speech Recognition Technology, where the device interprets the user has uttered. Speech recognition technology (speech-to-text) has grown more accurate and sophisticated in recent years. Previous generations of this technology enlisted a very narrow spectrum of comprehension; at each dialog step, they required application-specific grammar to define the wording of 'allowable' spoken responses. Such grammars could be described using the W3C GRXML standard (McGlashan et al., 2004). Newer generations of speech-to-text engines use general purpose language models that can be shared across applications and contexts. This, of course, represents a considerable step forward. However, there are still no open standards for language model description at the present time, which limits interoperability. -Empowering this process, Natural Language Understanding (NLU) capability - a branch of Artificial Intelligence - enables the comprehension of the meaning of a user’s human speech and clarification of underlying intentions. -The second category of conversational components - a Text-to-Speech System (TTS) - serves to convert language text into audible speech and can be implemented in software or hardware products. Normal language text may contain symbols, numbers and abbreviations that are converted to spoken words presented to the user via a speaker. The quality of a speech synthesizer is gauged by its similarity to the human voice and by its ability to be understood clearly. -The third category of conversational components, the Graphic User Interface (GUI) Adapter, comes into play in those unique instances where the user requests rich media content, such as buttons, images, and formatted text, to be visually displayed on the endpoint’s screen. The -1Standards exist for establishing audio connections between endpoints and platforms, for example in POTS telephone systems or VoIP connections. Smart speaker endpoints currently use interfaces that are specific to the platform. These may be published but they are not interoperable with other platforms. -Page 28 of 57 - -endpoint will generally only be able to display said content if it receives specially formatted input from the platform. Likewise, the platform will only be able to “understand” user GUI input if it is first “translated” into a form that platform can understand. -As defined by the World Wide Web Consortium (W3C) Voice Interaction Community Group (McGlashan et al., 2004), Dialog Manager is a component that receives semantic information derived from user input (via speech recognition + NLU), updates the dialog history, its internal state, then decides upon subsequent steps to continue a dialog and provides output. -The Dialog Manager gets its content and logic from the Web Content layer. The content might be hard coded or be dynamic via interactions with external APIs (for example it could output an address using a third-party address normalization API). -The Dialog Manager also passes information to the Session Manager and retrieves information from it. The Session Manager is a governance body within the voice architecture. It governs and administers conversational sessions with the user. This means controlling starting and stopping -points, identifying and verifying the speaker, verifying the conversational assistant, managing connectivity among multiple conversational assistants, and, significantly, maintaining the context necessary for an active conversation between a user and a bot/assistant. A unique session ID is generated for each session, allowing the context of the conversation to be maintained over time. -2.5 The Diversity of Today’s Conversational Assistance -The popularity of general-purpose consumer conversational assistance (Amazon Alexa, Google Home, Deutsche Telekom, Samsung Bixby, Baidu Xiaodu, and others) (Frost & Sullivan, 2022) has often obscured the growth of other parts of the conversational assistance ecosystem. -Firms such as Microsoft, Deutsche Telekom, Nuance (now a division of Microsoft), SoundHound, PeopleReign, Cerence, RASA, Mycroft.ai, Redfox.ai (and many others) (Frost & Sullivan, 2022) now offer enterprises the tools and services to develop conversational -Page 29 of 57 - -assistance solutions for devices, smart environments, personal transportation, customer service and employee support. -Enterprise investment in conversational AI worldwide is expected, by 2025, to reach a level double that of where it was four years prior; we’re now seeing per year annual growth rates of roughly 22-25 percent (Frost & Sullivan, 2022). -Analysts foresee the bulk of future growth in these areas: -● Customer service and employee support solutions – many of which have grown through the years from call center automation. These are often classified as interactive voice response (IVR) systems. -● Personal and autonomous transportation -● Smart environments, from homes to manufacturing facilities -● Hands-free environments, from surgical suites to warehouses - -SECTION THREE: AN APPROACH TO VOICE ASSISTANCE INTEROPERABILITY -3.0 Introduction -Section Three describes the Open Voice Network approach to the interoperability of voice assistants, and the current intent fulfillment patterns under investigation. -3.1. Modeling Interoperability on How People Communicate -The Open Voice Network acknowledges that a standard methodology for expressing and describing conversational interaction may be achievable -- but may not achieve the objectives of the ‘Jobs to Be Done’’ identified above, particularly for the Technical Innovators. -Attempts to describe (and standardize) how conversation can and should be modeled will quickly become outdated as the field of conversational interaction innovates. As an example, -Page 30 of 57 - -the latest generation of Voice System Infrastructure Providers and Technical Innovators put aside the widely adopted W3C VoiceXML and Speech Recognition Grammar standards in favor of creating new design paradigms based on less constrained speech-to-text and NLP technologies. -We assert this: human languages can be considered the ultimate interoperability standard. People have internal hidden motivations and knowledge. They use language to align their understanding of the world with tasks to be achieved. In some conversations people will have enormous levels of trust and shared understanding prior to the interaction. In others they will be complete strangers and may not trust each other. -We envisage that successful interoperability approaches will support a model of autonomous agents collaborating to achieve a goal together using standard ways to share information with each other, operating at various levels of trust and information sharing. -3.2 Interoperability: Mediation and Delegation -The modeling of voice assistant interoperability according to human communication leads us to identify two types of interaction: mediated and delegated. - Figure 3.2.a. Example of mediated and delegated communication. -Page 31 of 57 - -Figure 3.2.a shows two examples of how people inter-operate to achieve a conversational task, in this case between a 'user' and a helping ' assistant'. -In the top example of mediated communication, the user interacts with a single assistant. assistant A owns the conversation but may talk with others (e.g., assistant B) behind the scenes to achieve the goal. The client may be aware of the existence of assistant B but does not interact directly with them. Assistant A is the host of the communication. -The second example is one of delegated communication. To fulfill the intent of the user, assistant A passes control of the conversation to assistant B. Assistant A may monitor the ongoing interaction or may drop-out of the conversation altogether. In this example, assistant A is the initiator of the communication. -In this example of delegated communication, assistant A may need to talk behind the scenes with assistant B to establish whether assistant B is indeed the appropriate (or best) destination for delegation prior to their actual act of delegation. -To properly fulfill all user intents, assistant A will need to both mediate and delegate conversations. In some cases, the client may not need to communicate with assistant B resulting in a fully mediated conversation. For example, assistant A may place the user ‘on hold’ if this were a telephone conversation, or the dialog may be paused if it is an electronic channel such as chat. -In addition, assistant A may also be asked to serve as a delegate destination – to receive a request from assistant C to fulfill a user intent. -Analogous to this, Figure 3.2.b illustrates another aspect of interoperability between people. In situations where considerable trust exists between the collaborating assistants then brief referrals (in language) can be made along with the passing of a case-history. -Page 32 of 57 - - -A brief referral plus a detailed case history is passed to the next specialist. Each specialist interprets their role based on a referral request and the case history -Figure 3.2.b. Example of a delegated communication in which context (here, a cast history) is also shared. -The Open Voice Network proposes that, to the extent it is possible, models of inter-assistant and inter- assistant collaboration follow similar patterns to those followed by human assistants. -Key features of such a model would include: -● Each assistant has an identifiable identity and knowledge boundary. -● Each assistant will both mediate a communication (serving as a host assistant) and delegate a communication (serving as an initiating assistant) depending upon the user intent. assistants will also serve as a destination of a delegated communication. -● A desired scheme for interoperability will support collaboration as a mixture of mediated and delegated patterns. -● Delegation between assistants will be performed using brief requests. -● Levels of trust will vary between clients and assistants and this will affect the level of information sharing across boundaries. -● Where high levels of trust exist and there is shared understanding of how to represent knowledge, delegations and mediations will be supported by shared history and context. -Page 33 of 57 - -3.3 From the Human Model to Standards: OVON Interoperability Objectives -The OVON architecture seeks to establish standardized communication protocols between dialog assistants running on diverse platforms. -It does not, however, seek to standardize platform components or the format of content used to configure such components. OVON envisages a world in which conversational systems will continue to evolve in complexity and function. -As such the OVON architecture celebrates diversity in the following areas: -● The diversity of underlying technology (i.e., speech recognition, language understanding and dialog modeling) -● The diversity of endpoints [e.g., smart speakers, smart phone apps, dumb phones (POTS), audio-enabled web pages etc. -● The diversity of conversational design paradigms -● The diversity of labels used to represent semantic content. -To enable mediated and delegated communication by conversational assistants within an ecosystem of general purpose, consumer-centric conversational assistants, and independent, purpose- and brand-specific conversational assistants and assistants, the Open Voice Network will specify the messaging protocols and a destination and location services system that will provide a standardized approach to the way in which: -● basic linguistic information is shared between dialog assistants -● immediate linguistic context and history is shared between dialog assistants ● control is handed between dialog assistants -● dialog assistants negotiate trust. -● a spoken assistant name can be used to find an associated assistant. -Page 34 of 57 - -3.4 Architectural Patterns Under Investigation -In response to the Open Voice Network’s Jobs To Be Done review, and informed by the vision above, the Open Voice Network Architecture Work Group identifies several architectural patterns that, when combined together, could enable seamless intent fulfillment for users of different conversational assistants and assistants. -The architectural patterns that are being explored are: -● Dialog Delegation and Give-Back (additional detail to follow in Section 3.4.1) -A standard, distributed approach to the delegation and take-back of tasks from one platform to another. This is modeled on the idea of intercommunicating human assistants. -● Delegation Request and Conversation Context (Section 3.4.2) -A standard, extensible method of passing context between assistants. This is closely related to the delegation and take-back mechanism. This is modeled on the idea of maintaining case-history during sustained interactions. -● Dialog Interaction Payload (Section 3.4.3) -A standard, extensible format for the passing of linguistic events between components. This is expected to consolidate work done over many years by the linguistic and dialog engineering community. -● Dialog Component Interfaces (Section 3.4.4) -Standard component interfaces, built on top of the linguistic event format, starting with dialog interaction managers. This is expected to align as closely as possible with existing proprietary approaches. -● Discovery and Location (Section 3.4.5) -Standard patterns and mechanisms to allow assistants to publish and discover services from one another. -● Sharing and Protection of Data (Section 3.4.6) -A strategy for sharing data elements among conversational assistants. It also postulates a constraint enforcement mechanism that enforces privacy -constraints on shared data. -Page 35 of 57 - -These six different design patterns could be used separately or in conjunction with each other. They could also be adopted as part of other interoperability initiatives such as the Stanford University Open Voice Assistant Lab (OVAL) model discussed later in this paper. -Additional design patterns are in review by the Open Voice Network Architecture Work Group. Initial work on those patterns can be found in the Appendix, Section 1. -3.4.1. Dialog Delegation and Give-Back - -Figure 3.4.1 Delegation request -Figure 3.4.1 illustrates a typical delegation request. The user’s request is transferred from the user to a conversational assistant (which we will call the host) which examines the request. If the host determines that it cannot process the request itself, then the host negotiates (the dotted line in figure 3.4.1) with a second conversational assistant (the delegate) for possible processing. -A high-level acceptance/rejection mechanism delegates the request to the second conversational assistant (delegate conversational assistant). Some establishing or authenticating information will typically accompany the request to provide context about the initiating assistant and requested task. -Through this mechanism, the delegate conversational assistant will assume control, and then fulfill the user’s intent. In the most basic version of this scenario, the delegate assistant will then declare the task as finished. It may also ask the host assistant to reassume control. -Page 36 of 57 - -3.4.2 Delegation Request and Conversation Context -When, to fulfill a user intent, an initiating assistant seeks to delegate the conversation, the initiating assistant will frame its request so that all potential destination assistants can understand it. An establishing context for the dialog may also be necessary. -A delegation request from an initiating assistant requires these capabilities: -● Request: what the host assistant would like the potential delegate assistant to do: the user request -● Context: conversational history, plus other parameters (TBD). -3.4.3 Dialog Interaction Payload -The request payload contains information in standard formats that OVON will define. It consists of: ● Current request -● Conversational History -● Other data -The process of delegation may begin with an initiating assistant ( assistant X) passing part of all of the user utterance as its ‘request’ to a Destination and Location Service or directly to a destination assistant ( assistant Y). For example, one might say, “assistant X…” (wake word) “. . . “. . send me to assistant Y . .” ‘. . so that I might accomplish x, y and z”. -The potential destination assistant Y will determine if and how it can respond to the request. -The Architecture Working Group is exploring the use of a standardized dialog ‘payload’ format to support the passing of requests in a standard form. -3.4.4 Dialog Component Interfaces -The OVON Architecture Working Group identifies three levels that can be used to communicate a request from one assistant to another. -Page 37 of 57 - -● Level 0. An audio stream or file containing a natural language request. ● Level 1. A text string containing a natural language request -● Level 2. A structured semantic representation of a natural language request. -There is no architectural requirement for the delegation request to be built directly from the user’s language. The host delegate could synthesize a delegation request in its own natural language. When level 2 communication is used, both conversational assistants must have a common understanding of the semantic representations. In addition, we do not anticipate a requirement for a delegation request to be built directly from the initiating assistant’s language. The initiating delegate could synthesize a delegation request in its own language – natural or internal. This would also be of value in situations in which multiple natural languages are present, with either a multilingual human user or a multilingual assistant. The mechanisms for handling multilingual delegation are not yet specified. -In addition to the three layers noted above (audio, text, and semantic representations of speech) the following may also be represented in a conversational history: -● a user ID, standard for Identifying users -● a platform ID, standard for identifying the platform -● a conversational assistant, standard for identifying the conversational assistant ● a session ID, standard for identifying session -● inflection, standard symbol sets for annotating stress and intonation of speech ● affect, standard symbol sets for coding of emotions or affect of the speaker ● a dialog act, standard symbol sets for representing dialog acts (e.g., asking a question). -Other richer semantic representation formats will likely emerge over time. The set of supported layers will grow over time as proprietary formats are proposed for adoption. -The Open Voice Network may propose a standard format for each of these layers. At present, we envision the identification of specific layers as either mandatory or optional. -Regarding the Open Voice Network approach to standardization, the development of each of these layer schemas will draw heavily from formalized or de facto industry standards. These could include formal standards (e.g., TEI, UTF, W3C Pronunciation representation, etc.) or broadly adopted specifications from industry leaders. -Page 38 of 57 - -Each layer may possibly also be expressed using the Extended MultiModal Annotation Language EMMA (Johnson, 2009). -In an early proof of concept demonstration, the OVON Architecture Work Group successfully transmitted level 1 (text) messages between Mycroft, Magenta, and Genie conversational assistants. [MOAD summary.docx - Google Docs]. -3.4.4.1 An Illustrated Example: Pat Goes Shopping -Let’s return to the illustration of shopping with Pat, as presented earlier in section 1.5. Below, figure 3.3.4 breaks down the requests of Pat’s dialog with two conversational assistants, Shop (a home shopping assistant) and Payperson (a payment assistant). Note the three layers of interfaces and their interaction, as depicted here. - Figure 3.3.4: Current request and a subset of the conversational history of a Pat shopping dialog. -● Layer 0: Audio. An audio stream or file containing a natural language request. Voice is represented in computers as a sequence of bits, often called a wav file. Above, we show electronic representations as wave icons. -Page 39 of 57 - -● Layer 1: Text. A text string containing a natural language request. Above, the red icon, representing Pat’s utterance, is converted into the text: “Put two apples in cart.” -● Layer 2: The structured semantic representation of a natural language request, here shortened to "goals." Semantics (meaning) of a user goal is often represented using intent and slot container structures, which represent goals/actions to be performed and the parameters defining the actions. Above, Pat’s text expression is converted into the Intent, PutInCart, with the parameters “Apples” as the value of Food Item and “1” as the value of Quantity. -To complete the example, the dialog manager generates the text message, “Two apples were placed in cart,” which is converted by the TTS into a wav file (depicted by the green waveform graphic above) and then presented to the user. -. -3.4.5. Discovery and Location -An open, worldwide voice web will allow users to fulfill an intent through a mediated or delegated connection to any content source or conversational assistant, regardless of platform parentage. We believe that conversational assistance can and should work like the web. -Users will need to both find (e.g., the connection to a specific, named destination) and discover (e.g., the exploration of information and destination options within a topic or category) content in the worldwide voice web. -To meet this need, the Open Voice Network has initiated research into the development of what may be termed Discovery and Location Services. The services may include: -● a DNS-like service for the identification of available conversational assistants (by name and address) -● A standardized approach to metadata representation and organization -● the use of search engines, browsers, and aggregators that examine meta information about conversational assistants that provide the conversational assistant’s address -● the use of natural language requests sent from one assistant to another which the receiving assistant can use to decide if it can fulfill the request. -Page 40 of 57 - -A request to a Discovery and Location service from an initiating assistant will likely include text with the user’s request and contextual information. -The response from Discovery and Location services(s) and an identified destination assistant will include a digital address, configuration information, and privacy/security constraints. - -Figure 3.3.5: Schematic of the role of an envisioned Discovery and Location Service. -The schematic above illustrates the role of an envisioned Discovery and Location Service in the process of delegation. In a future worldwide voice web with millions, if not billions of destinations, a Discovery and Location Service (or Services) would enable an initiating assistant to find and connect with a destination assistant. In practice, it would determine which conversational assistants should be invited to the delegation and identify the digital addresses of the potential destinations. -3.4.6. Sharing and Protection of Data -An open, standards-based worldwide voice web of mediated and delegated dialogs will require a new and systemic approach to data protection. The Open Voice Network seeks to enable secure end-to-end connection between users and conversational assistants. This may include the encryption of messages, authentication of users as they initiate and progress through delegated dialogs, and authentication of conversational assistants in both mediated and delegated dialogs. -Page 41 of 57 - -Work in these areas is underway. -Earlier this year, the Privacy and Security Work Group of the Open Voice Network Technical Committee published for public review the guiding white paper Privacy Principles and Capabilities Unique to Voice. -From a perspective of a human user of voice assistance – and anticipating an OVON assertion of both mediated and delegated interoperability – the Privacy Principles white paper reviews the current landscape of regional and national privacy regulation and legislation, identifies primary human user risks and potential harms, and details four principles that the OVON believes should guide the development and implementation of voice assistance: -● Transparency: proactive communication – through an easily accessible and readily available interface -- to the individual user of data collection practices, general data usage, and data sharing policies. -● Consent: the expectation that the individual user must be allowed to give explicit, unambiguous agreement to the collection and processing of personal data. -● Limited Collection and Use: a restriction of the collection and analysis of raw and processed voice data – beyond that necessary for immediate dialog functionality – to stated and creator consent-given purpose. -● Control: the ability of the individual user to easily access, rectify, suppress, limit, oppose, and transport the data created by the individual user. -Under development now by the Privacy and Security Work Group are two additional papers: -● Data Security Specific to Voice, a review of current data security issues within voice assistance, and of data security issues inherent within multi- assistant voice interoperability. This paper is scheduled for publication in August 2022. -● Voice Assistance and the Privacy of Organizational Data, an exploration of why, when, and how organizations of all types (and especially enterprises) may protect proprietary data from undesired sharing and platform collection in both mediated and delegated interoperability. A first draft of this paper is scheduled for publication in the fourth calendar quarter of 2022. -Page 42 of 57 - -In parallel, the Ethical Use Task Force of the Open Voice Network published for public review the white paper Ethical Guidelines for Voice Experiences. -From the perspective of a human user of voice assistance, the Ethical Guidelines white paper identified voice-specific issues of ethical and moral concern, discussed rights that must be respected and values to be promoted, and addressed preventative measures that could be taken to protect users from voice-specific harm. -The paper also outlines a voice-specific ethical framework of five principles: -● Compliance: the acknowledgement of, and adherence to, ethical principles, standards, guidelines, and existing laws and regulation. -● Transparency: open and clear communication – in an easily accessible, understandable, and explainable user interface – regarding the collection, usage, and sharing of user and user-created data. -● Privacy Protection: not only adherence to the regulation and legislation that governs personal data of all types (General Data Protection Regulation (GDPR) – Official Legal Text, 2019) - especially voice data (Frost & Sullivan, 2022) - but the prioritization of data security and the vetting of the third parties that may, even with transparent consent, handle and process voice data. -● Inclusivity: working at all times to allow all to be heard – across languages, dialects, genders, ages, ethnicities, types and levels of disability. -● Accountability: maintenance of, and adherence to, highest ethical standards throughout the voice development, implementation, and operational value chain. -A next step for the Open Voice Interoperability Initiative (see Section Six, below) will be to translate the principles of these four papers into tangible, implementable technical guidelines and specifications. -The Work in Progress appendix to this paper (see section A.16, p. 69) contains an early proposal for an enforcement mechanism for privacy constraints. The approach consists of two mechanisms: (a) a mechanism to determine if data should be shared and with whom it is shared, and (b) a mechanism to monitor and control data after it is shared. -Page 43 of 57 - -SECTION FOUR: LESSONS FROM OTHER INTEROPERABILITY INITIATIVES -4.0 Introduction -As noted in Section One, the coming world of voice will be one of diversity, one marked by a multitude of voice assistants and agents. At present, however, the realm of general-purpose consumer-facing conversational assistants is one of singular, proprietary platforms that do not interact with one another. -As noted, the Open Voice Network envisions a future in which proprietary walls to content and connection are lowered, and users may seamlessly move from one assistant to another according to intent – a future in which voice operates broadly like the web, and not like apps on a mobile platform. -This section provides an overview of two important interoperability initiatives outside the Open Voice Network. We applaud –and are learning from – both. We also anticipate some level of adoption and adaptation of the concepts described here, as our work progresses. -4.1. Amazon Voice Interoperability Initiative (VII)™ -The Amazon Voice Interoperability Initiative (VII) enables the deployment of multiple conversational assistants on a singular endpoint device. Assessed with the Open Voice Network framework, it offers limited delegation – e.g., it does not allow the delegation of a conversation, but allows the user to leave one assistant and engage with another. -In a VII implementation, each assistant is activated via its own ‘wake word,’ enabling users to talk to the assistant of their choice in a secure manner by simply saying its name. For example, a user can interact with a specialized conversational assistant (such as Fridge) for specific refrigerator interactions, and leave the specialized conversational assistant and switch to a general-purpose conversational assistant (Amazon Alexa) for general purpose interactions. In -Page 44 of 57 - -this example, the Fridge conversational assistant would be designed with the vocabulary and skills specific to refrigerators and would bypass the vocabulary and skills that Alexa understands. Likewise, Alexa would not need to deal with the vocabulary and skills for the refrigerator. -The assistants interoperate with each other via the Multiagent experience (MAX) toolkit. The toolkit provides the MAX Library and a Sample Application that demonstrates the interoperability of Alexa and a second independent voice assistant. The MAX Library facilitates interoperability between voice assistants according to the guidance provided by the Voice Interoperability Initiative (VII) Multi-Agent Design Guide. - -Figure 5.1.1: Amazon Voice Interoperability Initiative -Within the VII guidelines, an assistant can transfer a user to a second assistant when it cannot directly fulfill a user request and is aware that the second assistant on the device can likely fulfill that request. No data or context is passed between assistants during a transfer, and the customer repeats their request directly to the second assistant without needing to say the wake word. -In addition, assistants interact with the device via Universal Device Commands. UDCs are commands and controls that a customer may use with any compatible assistant to control certain device functions, even if the assistant was not used to initiate the function. Examples -Page 45 of 57 - -include changing the volume of the device’s output speaker or stopping a sounding timer or music initiated from assistant B when assistant A is in control. -Comparison to the Open Voice Network proposal -● The focus of Amazon’s VII is on multi- assistant devices, i.e. several assistants accessed from one device. -● Delegation is limited to the assistants registered to the device. The Open Voice Network envisions a limitless number of potential voice-enabled destinations. -● Interoperability and assistant “discovery” is handled by the MAX toolkit and is limited to assistants registered and running on the device. In contrast, OVON allows for interactions between assistants potentially running exclusively on the cloud. -● In VII, assistants are discouraged to talk directly or pass information and context to each other. The user will need to repeat the queries to the second assistant. -● VII demands the presence of two running middlewares to mediate assistant interactions (the process running the MAX library and the device application and UDCs). -4.2. The Stanford Open Voice Assistant Laboratory (OVAL) Model -The Stanford University Open Voice Assistant Laboratory, under the direction of Professor of Computer Science Dr. Monica Lam, has proposed an open-source model for voice interoperability that enables mediation and partial delegation to multiple independent devices and services. (Lam et al., 2021). (Full disclosure: Dr. Lam is a valued advisor to the Open Voice Network, and the OVON is a financial supporter of the OVAL.) - -The Stanford OVAL model also follows the “Standardized Open Single Platform” model. The Stanford OVAL model gives developers the ability to collaboratively create “articles” through a device and services knowledge base known as “Thingpedia.” This empowers users to move from an initiating assistant to voice-enabled devices (i.e., a smart light bulb or smart factory sensor) and various services (a smart home system, a restaurant or retail website, or media properties such as Twitter or radio content). OVAL has demonstrated this model with several -Page 46 of 57 - -partners, using the OVAL “Genie” open-source conversational assistant as both mediator and initiator. -A strength of the Stanford OVAL model is that it allows the chaining and management of different articles combined in a single utterance (i.e., “When weather goes over 100F tweet out ‘oooh it’s hot’”). This also allows users to access multiple devices/services in a single request, although interoperability is restricted to the destinations, services, and use cases identified in Thingpedia. -Another significant strength of the OVAL model is that it enables users to enroll and save their devices/services. This allows users to interact with these devices/services securely and privately (without the need to authenticate separately and repeatedly). - -Page 47 of 57 - -Figure 5.2.1 Stanford OVAL architecture -The three core elements of the OVAL model: -ThingPedia -● ThingPedia is the knowledge base of what can be done by each device/service. -ThingSystem -● ThingSystem stores all device/service credentials for individual users allowing the platform to maintain privacy and security for the user. It runs the ThingTalk code that is provided by ThingPedia. -ThingTalk -● ThingTalk is the programming language that lies at the heart of OVAL platform. It gives the platform the power to connect IoT devices, web services, and database queries that are specified in ThingPedia. -Observations: -1. We perceive the OVAL model to be primarily one of single- assistant mediation. -2. If the OVAL model (with knowledge base Thingpedia and programming language ThingTalk) were widely adopted, it would deliver the capabilities of the Discovery and Location services identified above as a requisite for interoperability. -3. As the ThingPedia service expands, we fear that its ability to deliver discoverability may become strained. Currently Thingpedia has several hundred distinct “skills.” We must learn plans (UX, governance, etc.) for the scaling of the knowledge base to incorporate millions (or billions) of future skills. -Page 48 of 57 - -SECTION FIVE: FURTHER STUDY AND NEXT STEPS -This paper is but the first step on an important but challenging journey toward mediated and delegated voice assistance interoperability. -Visible next steps for the Open Voice Network Architecture Work Group include the following: -● The review, discussion, revision, and confirmation of the concepts shared in the Work in-Progress content in the Appendix to this white paper. -● The identification and assertion of messaging protocols that will allow independent voice assistants to mediate, delegate and serve as destinations for dialogs. With the encouragement of the Open Voice Network Steering Committee, the Architecture Work Group will first pursue adopting or adapting both existing standards and technologies that are now broadly used within the voice industry. -● Continued research as to mediated and delegated interoperability with existing and emerging voice and conversational AI services, including: -a) interactive Voice Response (IVR) systems inside corporate data security firewalls b) voice-enabled conversational bots -c) voice-enabled web applications -d) conversational AI implementations within enterprise software and processes e) enterprise-focused Conversational AI platforms. -● The development, with the OVON Voice Registry Work Group of the Technical Committee, of a standards-based approach to discoverability, findability, and location services. We anticipate publication of a separate document on this issue in the months ahead. -● The expansion of OVON messaging protocols to enable the sharing of multi-modal content. -Page 49 of 57 - -SECTION SIX: OPERATIVE VOCABULARY -Component: an identifiable part of a voice assistant or agent. A component provides a particular function or group of related functions. -Context: Information extracted from n prior utterances of the current conversation. This could include some or all of the following: information that has been inputted to outputted from, or inferred in Conversational Processors, and the information state of the Dialog Manager. Also known as ‘Conversational Context.’ -Conversation: a joint activity in which two or more assistants (human or automated) use linguistic forms and non-verbal signals (i.e., gestures) to communicate to achieve an outcome that meets a shared goal. -Conversation Event: a conversation event signals shifts in the conversation that may be acted upon. Such an event may occur at the beginning or ending of a Conversational Session, completion of a Conversation Processor, decoding of Conversation Information, changes to the state of a Conversation Endpoint, or changes to the status of a Conversation Stream. Any component with access to the system is allowed to generate a Conversation Event. -Conversation Facilitator: a component that coordinates communication between two or more Dialogue Systems and/or Processors during the course of one or more Sessions. This allows dialogue Systems and associated Processors to collaborate regardless of technology being used. Examples of Conversation Information include semantic, lexical, syntactic, and prosodic features. -Conversation Information Layer: an abstraction of a type of information in a Dialog System. A layer may be a specific type of acoustic, linguistic, non-linguistic, or paralinguistic features. Examples of layers would be Cepstral features, Phonemes, Intonation Boundaries, Words, Phrases, Turn Boundaries, Syllabic Stress, Discourse Move Type and specific Semantic representation schemes. -Conversation Processors: conversation information is encoded and/or decoded by one or more Conversational Processors. Conversational Processors may also take as input the output from another Conversation Processor. A Conversation Processor may generate Conversation Events and Conversation Streams. Example conversational processors include Automatic Speech Recognizers (ASR), Natural Language Processors (NLP), Dialog Managers, Text to Speech Synthesizers (TTS), etc. -Conversation Session: a particular conversation that consists of two or more Conversation Streams (see below) generated by two or assistants through one or more Conversational -Page 50 of 57 - -Endpoints. Sessions may be persistent, but they will often have a start-point and an end-point in time determined by one of the assistants or some other external event. -Conversation Stream: each Conversational Endpoint generates one or more Conversation Streams based upon the capabilities of the Endpoint and the preferences of assistant. A Conversation Stream is associated with a particular assistant and may include any media type including text, audio, video, and application UI events. -conversational assistant: a digital participant in a conversation. This may be an application with a consistent persona, such as Amazon Alexa, Google Assistant, the Target Google Assistant Action, a Facebook Messenger chatbot, or an IVR system at a bank or a human. conversational assistant is a common term utilized in Dialogue System research and university level instruction; the term often is used to describe a human participant in a conversation. For clarity of reference, however, the OVON will use the term "user" to identify a human participant. (See "user" below.) -Conversational AI: the set of technologies to enable automated communication between computers and humans. This communication can be speech and text. Conversational AI recognizes speech and text, understands intent, decipher various languages, and responds where it mimics human conversation. In some cases, it is also known as Natural Language Processing. -Conversational Context -See ‘Context,’ above. -Conversational Delegation: the passing of dialog layers and control between one Conversational Assistant and another to fulfill a user intent. The first assistant in the delegation sequence is the initiating assistant; the second is the destination assistant. -Conversational Endpoint: assistants conduct conversations using conversational endpoints; these may be a phone, mobile device, voice speaker, personal computer, kiosk, or any other device that enables an assistant to participate in a conversation. Endpoints may be referred to elsewhere as a "device" or a "channel." -Conversational Information Packets: information that relates to a specific period of time. Packets form the input and output of Conversation. -Conversational Mediation: the hosting of a dialog by a Conversational Assistant. In conversational mediation, the host assistant may fulfill a user intent by itself; it may access third-party data sources through API calls; or, it may introduce to the user a third-party application that is resident on the platform of the host assistant. A mediating assistant does not cede control, nor access to the data within the conversation. -Page 51 of 57 - -Conversational Platform: A group of technologies that are used as a base for one or more conversational assistants; also (see "Platform" below) a business model that harnesses and creates a large, scalable network of users and resources that can be accessed on demand. -Data: (per the Cambridge Dictionary): information, especially facts or numbers, collected to be examined and considered and used to help decision-making, or information in an electronic form that can be stored and used by a computer. Context (see above) is a subset type of the data accessed and used by the voice assistant system. -Dialog Manager (DM): handles the dynamic response of the conversation. It provides a more personalized response based on the action provided by the NLP to send back to the user. -Disambiguate: when the conversational platform hypothesizes two or more possible resolutions to a user utterance, it may ask the user for additional clarification or choose between the various interpretations to decide the user's correct intention. -Entity: a custom level data type and considered a concrete value to associate a word(s) in a query. It is a part of the structural machine translation. Also known as annotations. -Explicit Invocation: an invocation type where the user invokes the channel, and explicitly states a direct command to accomplish a specific task. The direct authority is to communicate directly to a registered voice application. -Implicit Invocation: an invocation type where the user invokes the channel by describing the target destination rather than naming the target. -Invocation: a part of the construct of the user's utterance during a conversation with a channel. An invocation describes a specific function that the guest wants, and solicits a particular response. -Intent: the identified action that the machine interprets based on the user's query. It is a part of the structured machine translation. Also known as a classifier. -Jobs to be Done: An approach to learning what will cause a customer to hire or bring your product or service into their life. -Natural Language Processing (NLP): a service and a branch of Artificial Intelligence that helps computers communicate with humans in their language and scales other language-related tasks. NLP helps structure highly complicated, unstructured human utterance and vice-versa. Natural Language Understanding is a subset of NLP that is responsible for understanding the meaning of the user's utterance and classifying it into proper intents. - - -Organization: a group of individuals brought together for a specific purpose, including the creation, transaction, and delivery of products or services. Examples would include a for-profit business, a not-for-profit group, or a government agency. -Platform: The collection of components (the environment) needed to execute a voice application. Examples of platforms include the Amazon and Google products that execute voice applications. -Query: user’s word requesting for specific function and expecting a particular response. -Speech-To-Text (STT): conversion of a representation of an utterance from audio to text.Also known as Automatic Speech Recognition (ASR). -Text-To-Speech (TTS): conversion of a representation of an utterance from text to audio. Also known as Speech Synthesis. -Technical Resource: it can be a publisher/developer. It can be a representative of an entity or independent party. Their role is to create an actual listing of the voice application. -Utterance: spoken or typed phrases. -User: a person who interacts with channels. -SECTION SEVEN: ABOUT THE OPEN VOICE NETWORK -The Open Voice Network (OVON) is a non-profit industry association dedicated to the development of standards for voice assistance transparency, consent, limited collection, and control of voice data that will make using voice technology worthy of user trust. In any reality, virtual or otherwise, we believe personal privacy should be respected as the default. The Open Voice Network operates as an open-source community within The Linux Foundation. It is independently funded and governed with participation from more than 120 voice practitioners and enterprise leaders from 12 countries. -The Open Voice Network community’s work is open source. We seek inclusive input and like to share our insights. At present, our work is focused in four areas: - - -● Interoperability, defined as the ability for conversational assistants to share dialogs (and accompanying context, control, and privacy) -● Destination registration and management, the ability of users to confidently find a destination of choice through specific requests, and for the providers of goods and services to register a verbal “brand” — similar to the Domain Name System (DNS) of the internet -● Privacy, with voice-specific guidance for both the protection of individual user data and that of commercial users -● Security, with a focus on voice-specific threats and harms. - -Please see our 2022 papers and support the Open Voice Network by visiting openvoicenetwork.org. -About The Linux Foundation -Founded in 2000, The Linux Foundation is supported by more than 1,000 members and is the world’s leading home for collaboration on open-source software, open standards, open data, and open hardware. Linux Foundation’s projects are critical to the world’s infrastructure including Linux, Kubernetes, Node.js, and more. The Linux Foundation’s methodology focuses on leveraging best practices and addressing the needs of contributors, users, and solution providers to create sustainable models for open collaboration. For more information, please visit us at linuxfoundation.org. -The Linux Foundation has registered trademarks and uses trademarks. For a list of trademarks of The Linux Foundation, please see its trademark usage page: www.linuxfoundation.org/trademark-usage. Linux is a registered trademark of Linus Torvalds. - - -Acknowledgements -This paper is authored by the Open Voice Network, with special thanks to the Architecture Work Group of the Technical Committee. Contributors: David Attwater, serving as Senior Research Scientist; Oita Coleman; Dr., Deborah Dahl, Senior Editor; Bruce Epstein, Co Moderator of the Architecture Work Group; Pradeep Gopal; Vineet Hingorani; Olga Howard; Carl Jahn; Kiran Kadekoppa; Dr. Jim Larson, Co-Moderator of the Architecture Work Group; Tobias Martens; Dr. Yaser Martinez-Palenzuela; Nick Myers; Shyamala Prayaga, Co Moderator of the Architecture Work Group; Elizabeth Robins; Dr. Dirk Schnelle-Walka; Nathan Southern; Jon Stine; Vadim Tarasevic; John Trammell; Boris Volfson. -We are grateful for the ongoing support of the Steering Committee of the Open Voice Network: Joel Crabb, Chair; Mirko Saul, Vice-Chair; Ali Dalloul, Bernhard Hochstätter, Doug Rogers, and Christian Wuttke. We wish also to thank and recognize individuals whose guidance, encouragement, and founding vision made this effort possible: Mike McNamara, Dan Cundiff, Kristi Dank, Maria Brinas-Dobrowski and Jay Kline; Ryan Steelberg and Sean King; Dr. Monica Lam and Jimmy Garcia-Meza; Reghu Ram Thanumalayan; Xuedong Huang; Bradley Metrock; Birgit Popp, Ulrike Stiefelhagen, and Anna Leschanowsky; Lawrence Lin. - - -SECTION EIGHT: REFERENCE LIST -Christensen, C., Hall, T., Dillon, K., & Duncan, D. (2016). Know Your Customers’ Jobs to be Done. Harvard Business Review, 94(9), 54–62. -https://hbr.org/2016/09/know-your-customers-jobs-to-be-done -European Data Protection Board. (2021, July 7). Guidelines 02/2021 on Virtual Voice Assistants. Retrieved August 11, 2022, from https://edpb.europa.eu/edpb_en -Fielding, R. T. (2022, June 1). RFC 9110: HTTP Semantics. RFC Editor. Retrieved August 11, 2022, from https://www.rfc-editor.org/rfc/rfc9110.html -Frost & Sullivan. Opportunities in the Conversational AI Market (2022, August); frost.com. Retrieved August 8, 2022 in advance of web publication. -General Data Protection Regulation (GDPR) – Official Legal Text. (2019, September 2). General Data Protection Regulation (GDPR). Retrieved November 8, 2022, from https://gdpr-info.eu -Haas, M. (2020, July 16). Understanding Conversational AI Tech. Interactions.Com. Retrieved August 11, 2022, from https://www.interactions.com/blog/technology/conversational-ai technology/ -IEEE Standards Information Network/IEEE Press. (2000). The Authoritative Dictionary of IEEE Standards Terms (IEEE 100), Seventh Edition (7th ed.). Institute of Electrical and Electronics Engineers (IEEE). -Johnston, M., Baggia, P., Burnett, D., Carter, J., Dahl, D., McCobb, G., & Raggett, D. (2009, February 10). EMMA: Extensible MultiModal Annotation markup language. World Wide Web Consortium. Retrieved August 11, 2022, from https://www.w3.org/TR/emma/ -Lam, M., Landay, J., & Manning, C. (2021). Launching a World Wide Voice Web. Stanford University. - -McGlashan, S., Burnett, D., Carter, J., Danielsen, P., Ferrans, J., Hunt, A., Lucas, B., Porter, B., Rehor, K., & Tryphonas, S. (2004, March 16). Voice Extensible Markup Language (VoiceXML) Version 2.0. World Wide Web Consortium. Retrieved August 11, 2022, from http://www.w3.org/TR/2004/REC-voicexml20-20040316/ -∞ -2022.08.18/08:50 - diff --git a/specs/Interoperable Dialog Event Object Specification 1.0 b/specs/Interoperable Dialog Event Object Specification 1.0 new file mode 100644 index 0000000..21cb222 --- /dev/null +++ b/specs/Interoperable Dialog Event Object Specification 1.0 @@ -0,0 +1,516 @@ +2023.06.09 +Draft Version 1.0.0 +Interoperable Dialog Event Object Specification Version 1.0 +The Open Voice Network +Architecture Work Group of the Technical Committee +Editor-in-chief: David Attwater Contributors: Emmett Coin +Deborah Dahl Jim Larson + +June 9, 2023 + +TABLE OF CONTENTS +CHAPTER 0. SCOPE AND INTRODUCTION +0.1 Document Scope +0.2 Dialog Events +0.3 A Foundation for Further Specialization + +CHAPTER 1. SPECIFICATION +1.1 Representation +1.2 Dialog Event Object +1.3 Span +1.4 Feature Objects +1.5 Confidence, Linking and Stand-off Annotation +1.6 JSON Path and the substring() extension +1.7 Alternates +1.8 Nomenclature + +CHAPTER 2. SCHEMA +CHAPTER 3. REFERENCES +CHAPTER 4. GLOSSARY OF TERMS +CHAPTER 5. DECISION LOG + + +Chapter 0. Scope and Introduction +0.1 Document Scope +This document specifies the format for Open Voice Network (OVON) interoperable dialog events. The requirements for this specification are given in Interoperable Dialog Packet Requirements [10]. This specification is generic. As described in the requirements there are many different potential uses of a dialog event. + +0.2 Dialog Events +Interoperable conversational systems will need to be able to process linguistic input and generate linguistic output. The components within such systems also need to send and receive such output in a standardized way. +The purpose of a dialog event is to define a generic standardized data structure that can be used in any component of a dialog system to express a ‘language event’, that is to say, any features associated with a phrase, utterance or part of an utterance. Dialog events span a certain time period and are associated with a single speaker. +These events are used to represent user inputs and system outputs of various types, primarily linguistic inputs and outputs such as speech or text, but also potential multimodal inputs and outputs, such as selections on a touchscreen or images presented by a system. Dialog events can be used to express whole utterances, phrases, or other slices of time. +Components of dialog systems may receive an event and add features to it, for example, a natural language interpretation component may receive an event containing a text feature and add a semantic feature to it as an interpretation of the text. +It is anticipated that in the future, events could be joined together to form streams, for example, to represent continuous input or output of a speech-to-text engine. + +0.3 A Foundation for Further Specialization +This specification defines a generic format into which dialog features in different formats can be gathered together into a single event. It also describes how such features can reference each other. +It does not define which formats (mime types) should be present, how features should be named, or which features might be required under what circumstances. This will be left to additional specifications, which we term derived specifications, built upon this one. +For example, dialog events could be used to carry prompt and response information for a dialog system, annotate spoken dialog between two human speakers, or be used as part of an interface to a natural language component in a text processing solution. +We are anticipating that a future OVON standard for representing utterances in a dialog system or dialog history will be developed as one example of a derived specification. + +Chapter 1. Specification +1.1 Representation +Dialog event objects will be represented as a JSON [1] object in a string format. The JSON dialog event object is not intended to be a stand-alone document. It is intended to be included in a larger data structure as an object with a re-usable standard structure + +1.2 Dialog Event Object +Each dialog event object has a unique ID, is associated with a single speaker (person or machine), spans a period of time, and contains features representing different inter-related aspects of the event. + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT0.1045" + }, + "features": { + "my-audio-feature": { ... } + "my-text-feature": { ... } + + + "my-semantic-feature": { ... } + "my-custom-feature": { ... } +} } +Figure 1. Example of a dialog event container. +Figure 1 shows an example of a dialog event object. Each dialog event object has a unique id and relates to just one speaker identified by the speaker-id. The event can optionally be linked to a previous event object from the same speaker using the previous-id. The span object represents the time-span of the event. +The features: id, speaker-id, span, and features are mandatory. The previous-id object is optional. The features section will typically contain one or many feature objects but may be empty. Each feature object is identified with a key which can have any arbitrary feature name. As dictionary keys, feature names must be unique within a dialog event. + +1.3 Span +The span object describes a period of time during which the event occurred. It is this span that distinguishes the dialog event from a visual user-interface event which occurs at a point in time. Span objects must contain either an absolute start-time or a relative start-offset but not both. They can also optionally contain either an end-time or end-offset but not both. Absolute times are represented in ISO 8601 [2] format, more specifically the simpler interoperable variant specified in IETF RFC 3339 [3]. Relative durations are represented in ISO 8601 [2] duration format. +The top-level span object may contain a start-offset rather than an absolute start-time. This start-offset specifies the start time as the time elapsed since an absolute reference time is defined outside of the object. For example, the event might be contained in a dialog history object and this offset might be relative to the start-time of the conversation represented by that object. +As will be explained below, span objects can also be included in individual tokens within features of the dialog event. + + +1.4 Feature Objects + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT3.1045" +}, + “features”: { + "my-audio-feature": { + "mime-type": "audio/wav", + "tokens": [ + { + "value-url": "http://localhost/xyz1234.wav" +} ] +}, + "my-text-token-feature": { + "mime-type": "text/plain", + "lang": "en", + "encoding": "UTF-8", + "token-schema": "BertTokenizer.from_pretrained(bert-base-uncased)" + "tokens": [ + {"value": "what"}, + {"value": "is"}, + {"value": "the"}, + {"value": "weather"}, + {"value": "forecast"}, + {"value": "for"}, + {"value": "tomorrow"} +] } +} + + + } +Figure 2. A dialog event containing an audio feature object and a tokenized text feature object. As a minimum, each feature object must contain a mime-type and an array of token objects +named tokens. + +Figure 2 shows a simple example of a tokenized text feature. In this example, the mime-type is plain text and the feature is broken into a sequence of tokenized words. Each word is represented by a token object. +This is just one example of how text can be represented in a feature. The tokens array contains zero or more token objects, and each token object must contain either a value object or a value-url string. The value-url string references an external document that contains the value. The reference is a string in the format of a Universal Resource Locator (URL) as shown in the my-audio-feature in Figure 2. +The mime-type is mandatory and defines the format of the value object or the document referenced by the value-url. The format of mime-type is defined in IETF RFC-2231 [6] +The object lang defines the language of the feature according to IETF BCP 47 language Tag [4]. This value is optional and remains undefined if not included. This value informs how the feature should be interpreted. If the feature contains multilingual content it is recommended that this value remains undefined or is set to the dominant language of the feature. +The encoding object defines the text encoding within the value objects. Values should be one of "ISO-8859-1" (See ISO/IEC 8859-1 [11]) or "UTF-8" (See RFC3629 [12]). This value is optional and remains undefined if not included. +The token-schema object is included to allow reference to standard sets of token values or symbols for certain features with specific mime-types. For example, it could be used to specify which phonetic pronunciation alphabet is present in a pronunciation feature, or refer to a specific standard set of intents for a given domain in a semantic feature. The format of this object is not mandated and is expected to be defined in derived specifications. This value is optional and remains undefined if not included. Figure 2 shows one example of how the token-schema might be used to reference a specific text tokenizer. Figure 4 shows an example of + how token-schema might be used to reference a standard semantic schema. Both examples are illustrative, not normative. + .. +"my-text-token-feature": { +"mime-type": "text/plain", +"tokens": [ + { + "value": "what", + "span": { + "start-offset": "PT0.0210", + "end-offset": "PT0.1457" + } +}, { + "value": "is", + "span": { + "start-offset": "PT0.1460", + "end-offset": "PT0.1976" + } +}, .. { + "value": "tomorrow", + "span": { + "start-offset": "PT2.3082", + "end-offset": "PT3.0784" + } +}, ] +} .. + Figure 3. A feature object containing tokens objects with spans. +Each token object can also optionally contain a span object. This indicates the time span of this particular token. As described in section 1.3 Span, spans can be defined as absolute times or relative offsets. In a token object start-offset and end-offset are relative to the start-time of the parent dialog event (or the implied dialog event start time defined by its start-offset). +Figure 3 shows an example where the tokenized words of the text feature for Figure 2 are each given an individual span. + +1.5 Confidence, Linking, and Stand-Off Annotation +Each feature should represent just one aspect of the dialog event. Each token object within each feature can be linked to another token object in the event using the object links. + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT0.1045" +}, + "features": { + "my-text-feature": { + "mime-type": "text/plain", + "lang": "en", + "encoding": "UTF-8", + "tokens": [ + { + "value": "what is the weather forecast for tomorrow", + "confidence": 0.99 +} ] +}, + "my-semantic-feature": { + "mime-type": "text/x.my-semantic-mime-type", + "token-schema": "standard-intents.org/schemas/weather.xml", + "tokens": [ +{ +"value": { + 9 + + "name": "intent", + "label": "WeatherForecast" + }, + "links": ["$.my-text-feature.tokens[0].value"], + "confidence": 1.0 + }, +{ +"value": { + "name": "date", + "label": "02-14-2023" + }, + "links": [ +"$.my-text-feature.tokens[0].value.substring(34,41)" ], + "confidence": 0.87 + } +] } +} } +Figure 4. A simple example of linking between features with confidence assigned to a feature. +Figure 4 shows an example of a semantic feature which links to the text that it represents. +The optional links object is an array of references to the value (or parts of the value) of other token objects using a JSON Path with some extensions. There is no restriction on how many links a feature can have. +Unlike the text in Figure 2, in this example, the text is not tokenized. The semantic token with the name "intent" is linked to the whole text of the utterance as defined by the JSON path $.my-text-feature.tokens[0].value. The semantic token with the name "date" however is linked to a sub-string of the source text - the word ‘tomorrow’ - via the JSON Path $.my-text-feature.tokens[0].value.substring(34,41). Section 1.6 describes how to use JSON Path references in more detail. +The optional confidence object is a number carrying a measure of the confidence that the information contained in the associated value is ‘correct’. Confidence values are expressed as a probability (real number between 0.0 and 1.0). Derivative specifications may override this range for specialized situations. +Recall that the format of a given value object is defined by the mime-type and will vary between features. Feature layering and cross-referencing are both important parts of the dialog event standard. Together they permit the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. This approach is termed stand-off-annotation. [9] + +1.6 JSON Path and the substring() extension +Individual elements of the linked object are defined as strings containing JSON Paths. JSON Path is a notation for querying and manipulating data stored in JavaScript Object Notation (JSON). +At the time of publication, there is not an agreed standard definition for JSON Path. There are a number of implementations based on JSON Path [7] that are broadly compatible. There is also an IETF task force RFC6901 [8] working on a common standard. +JSON Path expressions in links objects are cross-references from one token object to the value, or part of a value, of another token object. For these references, the root document (designated by ‘$’) is the features object in the current dialog event. There is currently no way to specify a link between features that are not contained in the same dialog event. +Implementations should support one additional non-standard extension to JSON Path - namely the substring() function. This extension is added to allow token objects to reference a specific substring within another token object, for example, a word, phrase, or arbitrary sequence of characters. +The substring function can be invoked at the tail end of a path. The input to this function is the output of the preceding path expression. The substring function takes two ordered comma-separated parameters, the index of the start and end character in the substring. The first character of the string is at the index zero. +For example the JSON Path: $.my-text-feature.tokens[0].value.substring(34,41) will return the 34th to the 41st character of string contained within the object referred to by the JSON Path $.my-text-feature.tokens[0].value. If the referred object is not a string, or the substring is out-of-bounds then the resulting behavior is undefined. + +Token objects can contain content that is specified externally via a value-url string. If this external content is expressed in JSON format then the JSON path should be interpreted as if the value had been specified locally. This is functionally equivalent to reading the content of the document referenced by value-url and placing it in the equivalent local value object before applying the links reference. + +1.7 Alternates + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00" + }, + "features": { + "my-text-feature": { + "mime-type": "text/plain", + "tokens": [ + { + "value": "what is the weather forecast for tomorrow", + "confidence": 0.96 +} ], + "alternates": [ + [ + { + "value": "what is the weather forecast for thursday", + "confidence": 0.73 +} } +] } +} } + 12 + +Figure 5. Defining alternate values for a feature. +Spoken or written language rarely has a single unambiguous interpretation. For this reason, the dialog event object allows for multiple interpretations of a given event. As has already been seen, the tokens object is used to represent the preferred interpretation of a given feature. In addition to this, the optional alternates array can be used to express alternate interpretations of this feature. For example, in Figure 5 it can be seen that the preferred interpretation of the feature named my-text-feature has the value “what is the weather forecast for tomorrow” and a confidence of 0.96. The alternates array contains one additional interpretation of the utterance with the value “what is the weather forecast for thursday” and a confidence of 0.73. +The alternates object is an array of an array. Each object in the outer array represents an alternative interpretation and each inner array represents the token objects for this interpretation. These inner token objects have the same format as array elements in the tokens object. +The outer array can contain zero or more interpretations. An empty alternates array implies that there are no alternate interpretations and is equivalent to the alternates object not being present in the event. + +1.8 Nomenclature +This specification uses ‘kebab-case’ (i.e. hyphenated lowercase) for all nominal property names, for example, start-offset and value-url. User-defined object keys such as the dictionary names used in the features section can use any valid JSON expression. It is recommended but not mandated that any normative key names in derived specifications follow the same kebab-case convention. + + +Chapter 2. Schema +The structure of a JSON dialog event is defined below as a JSON Schema. + { + "$id": "https://openvoicenetwork.org/schema/dialog-event.json", + "$schema": "https://openvoicenetwork.org/schema", + "description": "A representation of a 'language event’ that is to say +any information associated with a phrase, utterance or part of an +utterance.", + "type": "object", + "required": ["id", "speaker-id", "span", "features" ], + "properties": { + "id": { + "type": "string" + }, + "previous-id": { + "type": "string" + }, + "speaker-id": { + "type": "string" + }, + "span": { "$ref": "#/$defs/span" }, + "features": { + "type": "object", + "patternProperties": { + ".*": { "$ref": "#/$defs/features" } + } +} }, + "$defs": { + "features": { + "type": "object", + "required": ["mime-type" , "tokens" ], + "properties": { + "encoding": { + "type": "string", + "description": "The text encoding of the token values" + }, + "mime-type": { + "type": "string", + "description": "The mime-type of the token values" +}, "lang": { + "type": "string", + "description": "The language of the token values" + }, + "token-schema": { + "type": "object", + "description": "A schema restricting the token values" + }, + "tokens" : { + "type": "array", + "items": { + "$ref": "#/$defs/token" + }, + "alternates": { + "type": "array", + "items": { + "type": "array", + "items": { + "type": "object", + "$ref": "#/$defs/token" + } +} } +} } +}, "token": { + "type": "object", + "anyOf": [ + { "required": ["value"] }, + { "required": ["value-url"] } + ], + 15 + + "properties": { + "value": { + "type": ["number","string","object","array","boolean"] + }, + "value-url": { + "$ref": "#/$defs/url" + }, + "confidence": { + "type": "number" + }, + "span": { "$ref": "#/$defs/span" }, + "links" : { + "type": "array", + "items": { + "$ref": "#/$defs/jsonpath" + } +} } +}, "url": { + "type": "string", + "description": "Any valid URL" +}, +"jsonpath": { + "type": "string", + "description": "an expression in JSON Path syntax" +}, "span": { + "anyOf": [ + { "required": ["start-time"] }, + { "required": ["start-offset"] } + ], + "properties": { + "start-time": { + "$ref": "#/$defs/iso-time" + }, + "end-time": { + "$ref": "#/$defs/iso-time" + 16 + + }, + "start-offset": { + "$ref": "#/$defs/iso-duration" + }, + "end-offset": { + "$ref": "#/$defs/iso-duration" +} } + }, + "iso-time": { + "description": "A string in ISO 8601 absolute format." + }, + "iso-duration": { + "description": "A string in ISO 8601 duration format." +} } +} + Chapter 3. References +[1] https://www.ecma-international.org/publications-and-standards/standards/ecma-404/ ECMA-404 The JSON data interchange syntax +[2] https://www.iso.org/iso-8601-date-and-time-format.html ISO 8601 Date and Time Format. +[3] https://datatracker.ietf.org/doc/html/rfc3339 Newman, Chris; Klyne, Graham (July 2002). Date and Time on the Internet: Timestamps. IETF. doi:10.17487/RFC3339. RFC 3339. Archived [4] https://www.rfc-editor.org/rfc/rfc5646.txt RFC 5646. BCP 47. Tags for Identifying Languages. [5] https://www.ietf.org/rfc/rfc4646.txt RFC 4646 Regarding Best Practice for Tags for Identifying Languages +[6] https://datatracker.ietf.org/doc/html/rfc2231 MIME Parameter Value and Encoded Word Extensions: Character Sets, Languages, and Continuations +[7] https://goessner.net/articles/JsonPath/ JSONPath - XPath for JSON. Stefan Goessner. +[8] https://www.rfc-editor.org/rfc/rfc6901 RFC6901. JavaScript Object Notation (JSON) Pointer [9] Pose-Rodriguez, Javier & Lopez, Patrice & Romany, Laurence. (2014). A Generic Formalism for Encoding Standoff annotations in TEI. +[10] https://openvoicenetwork.org/docs/2832-2/ Interoperable Dialog Packet Requirement Specification. +[11] https://www.iso.org/standard/28245.html ISO/IEC 8859-1:1998 Information technology — 8-bit single-byte coded graphic character sets — Part 1: Latin alphabet No. 1 +[12] https://datatracker.ietf.org/doc/html/rfc3629 RFC3629. UTF-8, a transformation format of ISO 10646 + +Chapter 4. Glossary of Terms +Term +Definition +dialog event +A linguistic event in a spoken or written monologue or dialog between two or more speakers. +dialog event object +A JSON object encoding a dialog event as part of an interface to a natural language component in a text or speech processing solution or dialog system. +dialog event feature +A layer of information of a certain type associated with the dialog event. +confidence +A number representing a measure of the confidence that the information contained in the associated value is ‘correct’. +token encoding +The specific encoding used to represent text in a token. +token link +A link from one token in a feature to part or all of another token in a feature used to implement stand-off annotation. +JSON path +An unambiguous reference to part of a JSON object. +language code +A code representing the language (e.g., American English, British English, New Norsk, etc.) +feature mime type +The type of token values in a specific dialog event feature. +dialog event object id +The unique identifier of the dialog event object +token span +Identifies the span of time for an individual token object +dialog event span +Identifies the span of time for the dialog event + +dialog event speaker id +A unique id of the human or machine associated with content of the dialog event +stand-off-annotation +A method of feature layering and cross-referencing that permits the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. +feature token +A representation of part of the information that makes up a feature value. +token object +A JSON object representing a feature token which defines the value of the feature and other associated information such as its span and how it links to other feature tokens. +derived specification +An derivative standard, built upon this specification, which defines a specific way to use a dialog event in a particular context. + +Chapter 5. Decision Log +This section documents some of the key design decisions that were made by the team during the development of this specification. It is informative, not normative. +Issue Topic +Issue and Decision +Nomenclature +Issue: What style of object name should we adopt in the dialog event? +Decision: We adopt ‘kebab-case’ (i.e. hyphenated lower case) for all property names, for example ‘start-offset’ and ‘value-url’. +links +Issue: Should we allow multiple links for a token object? +Decision: Yes so that one feature can reference more than one other feature. +Span start and end times +Issue: Should start times and end times be mandatory or optional? +Decision: Start times are required. End times should be optional because events containing text to be generated via + +Start and end offsets +Time Zones +Features +Speaker Roles +Representation +Representation of Alternates +text-to-speech cannot know when the generated text will actually end in real time. For such generative events the start time will be the time that the utterance is intended to start. It cannot anticipate actual delays in processing. We would anticipate that when the text is processed to generate the audio, the audio will have full span. At this point the end time and the start time may be modified to represent the viewpoint of the downstream process that generated it (i.e. the TTS may add an end time and also change the start time to the point where an event is re-transmitted). +Issue: Should start and end times just be absolute or should we also support relative options such as offsets? +Decision Times can be absolute or relative but not both. The optionality is not complex and supports common use cases. +Issue: Does it matter that ISO can represent different time zones and will it cause a problem if servers are in different time zones? +Decision: No. It shouldn’t matter as long at consumers of messages regularize for their own time zone before using the time. ISO 8601 carries the time zone information in the format. +Issue: What is the cardinality of features object? +Decision: We expect at least one feature but in the spec we will allow for zero or more features. +Issue: Should an event carry speaker roles (e.g. agent, customer, bot, etc.)? +Decision: No. Because the role of a speaker is defined by the consumer of the event not the event itself +Issue: Should we be syntax-neutral or mandate a specific representation? (e.g. YML, JSON and XML) +Decision: We will specify messages as JSON. It is straightforward to translate it at source and destination into other formats. +Should a feature object contain a simple list of lists of tokens or should we separate out alternates into a different object? + Decision: We will treat alternates separately to keep the object simple for systems that do not represent uncertainty in their events. +Feature Cross-Referencing +Issue: Should we use simple array indexing or something more sophisticated such as JSON Path? +Decision: We will use path indexing via JSON Path. Using a path notation means that we do NOT need to insist that values are arrays. They can be any kind of object. This will allow us to reference strings by character indexes OR reference arrays by array indexes OR reference dictionaries by their dictionary names. +External Links +Issue: How do we refer to features in external events? ( This was originally a requirement in the requirements specification.) +Decision: We leave this open for now. See comments on streaming. +Feature Span +How do we attach a time span to a feature? +Decision: We allow the span object to be added to token objects. + OVON Interoperable Event Restrictions + Confidence +Issue: Should we leave the range of confidence values open for derived specifications to decide? +Decision: No. We will mandate a floating point number between 0..1 and leave it up to derived specifications to override this. +Streaming +Issue: Is a stream represented by multiple events in the same stream? +Decision: The use of features and events for streaming needs to be worked through more carefully. It is possible that it will require each feature to have a separate ID and span to allow different spans for different features and to allow continuously evolving interpretations in for example an asynchronous dialog system. + + +Issue: Does streaming apply to both speech and text? Decision: Yes +JPath or JSON Path? +Should we adopt JPath or JSON Path? +Decision: Neither are true standards. JPath appears to be IBM specific. We will adopt JSON Path and extend it to support substring matching. We will also aim to adopt any emerging IETF standard if it is completed. We welcome comments on other standard ways to refer to JSON paths +encoding +Issue: Is it the best approach to follow IETF HTTP in allowing only two encoding types "ISO-8859-1" [ISO-8859-1] or "UTF-8" [RFC3629]. +Decision: Yes for now but we welcome comments. +Token array +Issue: This representation "tokens":[ +{"token": +{ +} }, +{"token": +{ +} } +"value": "what", "span": { +---etc } +"value": "what", "span": { +---etc } +] +is different than the unnamed array elements. This may be valid with JSON indexing. But it permits/forces a JSONtoXML conversion to decide what to "name" the subelements. Decision: to be discussed +Value of “tokens” +Issue: In 1.5 Confidence .... +The value refers to a complete string: +"tokens": [ { +"value": "what is the weather forecast for tomorrow", +"confidence": 0.99 } + ] +But in the earlier example they were single word "tokens". Will +developers need to scan the char-strings to know what they are getting? +Decision: to be discussed +Alternates +Issue: . In the "alternates" section +Figure 5 is missing "]" and if you wanted it to represent +multiple alterates then we should have another alternate array element in the example. +This gets verbose if the "tokens" array is single words as in figure 2. +Decision: to be discussed +Tokens +It looks like according to the Schema the things under "tokens" and "alternates" must be a "token", but the examples don't reflect that. +Decision: to be discussed + + From cc5399b7efbca5afa1fa2a06b2a4128fb15a2ddf Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:03:02 -0500 Subject: [PATCH 28/35] Update Interoperable Dialog Event Object Specification 1.0 --- specs/Interoperable Dialog Event Object Specification 1.0 | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/specs/Interoperable Dialog Event Object Specification 1.0 b/specs/Interoperable Dialog Event Object Specification 1.0 index 21cb222..b41c83a 100644 --- a/specs/Interoperable Dialog Event Object Specification 1.0 +++ b/specs/Interoperable Dialog Event Object Specification 1.0 @@ -1,3 +1,9 @@ +*********** + +--- +layout: default +title: Interoperable Dialog Event Object Specification Version 1.0 + 2023.06.09 Draft Version 1.0.0 Interoperable Dialog Event Object Specification Version 1.0 From 57ee451697bd5cefdeab9d16e56524dadc030b1a Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:06:39 -0500 Subject: [PATCH 29/35] Update Interoperable Dialog Event Object Specification 1.0 --- specs/Interoperable Dialog Event Object Specification 1.0 | 1 + 1 file changed, 1 insertion(+) diff --git a/specs/Interoperable Dialog Event Object Specification 1.0 b/specs/Interoperable Dialog Event Object Specification 1.0 index b41c83a..61ba9a4 100644 --- a/specs/Interoperable Dialog Event Object Specification 1.0 +++ b/specs/Interoperable Dialog Event Object Specification 1.0 @@ -3,6 +3,7 @@ --- layout: default title: Interoperable Dialog Event Object Specification Version 1.0 +parent: specs 2023.06.09 Draft Version 1.0.0 From 64d3246ba78d3198776444afe20379bad16359c9 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:10:20 -0500 Subject: [PATCH 30/35] Create test --- technical-committee/test | 34 ++++++++++++++++++++++++++++++++++ 1 file changed, 34 insertions(+) create mode 100644 technical-committee/test diff --git a/technical-committee/test b/technical-committee/test new file mode 100644 index 0000000..f36f4be --- /dev/null +++ b/technical-committee/test @@ -0,0 +1,34 @@ +*********** + +--- +layout: default +title: Meeting Notes +parent: Technical Committee +--- +# Notes of the Open Voice Network Technical Committee Meeting - June 9, 2023 +**The Meeting Began at 11:01am EDT.** + +**Attendees: J. Stine, N. Southern, O. Coleman, T. Martens, J. Larson, D. Dahl, B. Epstein, H. Pappas** + +**Basic Welcome to Meeting Attendees -- T. Martens** + +**Notice of Recording - T. Martens** + +**Reading of Linux Foundation Anti-Trust Statement - N. Southern** + +## Minutes Approval of May 12, 2013 Technical Committee Meeting - T. Martens ## + +*Mr. Stine put forth a first approval motion; Dr. Larson seconded that motion; Mr. Southern marked the minutes duly approved.* + +## Review of Agenda, Opening Comments - T. Martens ## + +*Mr. Martens noted the following agenda items for today's meeting* + +* Bürokratt update on Estonia - Working Plan - D. Dahl +* Interoperability Roadmap update - D. Dahl +* Interoperability Webinar on June 15th - D. Dahl +* Recruitment for Demonstrators- J. Stine & T. Martens +* Trustmark Initiative Update/Self-Assessment Tool - O. Coleman +* New Terms for OVON Glossary - J. Larson +* Miscellaneous Comments/Questions - Group +* Closing Remarks - T. Martens From 33b398b3a060b082d5ec533ba608aecf408737f2 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:15:18 -0500 Subject: [PATCH 31/35] Update Interoperable Dialog Event Object Specification 1.0 --- specs/Interoperable Dialog Event Object Specification 1.0 | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/specs/Interoperable Dialog Event Object Specification 1.0 b/specs/Interoperable Dialog Event Object Specification 1.0 index 61ba9a4..009b2ca 100644 --- a/specs/Interoperable Dialog Event Object Specification 1.0 +++ b/specs/Interoperable Dialog Event Object Specification 1.0 @@ -1,12 +1,11 @@ *********** - --- layout: default title: Interoperable Dialog Event Object Specification Version 1.0 parent: specs -2023.06.09 -Draft Version 1.0.0 +## 2023.06.09 +## Draft Version 1.0.0 Interoperable Dialog Event Object Specification Version 1.0 The Open Voice Network Architecture Work Group of the Technical Committee From aca2a51547f1cbdda9403fffe99763c2d66b9773 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:20:20 -0500 Subject: [PATCH 32/35] Create InteropDialogEventSpecs.md --- specs/InteropDialogEventSpecs.md | 521 +++++++++++++++++++++++++++++++ 1 file changed, 521 insertions(+) create mode 100644 specs/InteropDialogEventSpecs.md diff --git a/specs/InteropDialogEventSpecs.md b/specs/InteropDialogEventSpecs.md new file mode 100644 index 0000000..eda6d60 --- /dev/null +++ b/specs/InteropDialogEventSpecs.md @@ -0,0 +1,521 @@ +*********** +--- +layout: default +title: Interoperable Dialog Event Object Specification Version 1.0 +parent: specs + +## 2023.06.09 +## Draft Version 1.0.0 +Interoperable Dialog Event Object Specification Version 1.0 +The Open Voice Network +Architecture Work Group of the Technical Committee +Editor-in-chief: David Attwater Contributors: Emmett Coin +Deborah Dahl Jim Larson + +June 9, 2023 + +TABLE OF CONTENTS +CHAPTER 0. SCOPE AND INTRODUCTION +0.1 Document Scope +0.2 Dialog Events +0.3 A Foundation for Further Specialization + +CHAPTER 1. SPECIFICATION +1.1 Representation +1.2 Dialog Event Object +1.3 Span +1.4 Feature Objects +1.5 Confidence, Linking and Stand-off Annotation +1.6 JSON Path and the substring() extension +1.7 Alternates +1.8 Nomenclature + +CHAPTER 2. SCHEMA +CHAPTER 3. REFERENCES +CHAPTER 4. GLOSSARY OF TERMS +CHAPTER 5. DECISION LOG + + +Chapter 0. Scope and Introduction +0.1 Document Scope +This document specifies the format for Open Voice Network (OVON) interoperable dialog events. The requirements for this specification are given in Interoperable Dialog Packet Requirements [10]. This specification is generic. As described in the requirements there are many different potential uses of a dialog event. + +0.2 Dialog Events +Interoperable conversational systems will need to be able to process linguistic input and generate linguistic output. The components within such systems also need to send and receive such output in a standardized way. +The purpose of a dialog event is to define a generic standardized data structure that can be used in any component of a dialog system to express a ‘language event’, that is to say, any features associated with a phrase, utterance or part of an utterance. Dialog events span a certain time period and are associated with a single speaker. +These events are used to represent user inputs and system outputs of various types, primarily linguistic inputs and outputs such as speech or text, but also potential multimodal inputs and outputs, such as selections on a touchscreen or images presented by a system. Dialog events can be used to express whole utterances, phrases, or other slices of time. +Components of dialog systems may receive an event and add features to it, for example, a natural language interpretation component may receive an event containing a text feature and add a semantic feature to it as an interpretation of the text. +It is anticipated that in the future, events could be joined together to form streams, for example, to represent continuous input or output of a speech-to-text engine. + +0.3 A Foundation for Further Specialization +This specification defines a generic format into which dialog features in different formats can be gathered together into a single event. It also describes how such features can reference each other. +It does not define which formats (mime types) should be present, how features should be named, or which features might be required under what circumstances. This will be left to additional specifications, which we term derived specifications, built upon this one. +For example, dialog events could be used to carry prompt and response information for a dialog system, annotate spoken dialog between two human speakers, or be used as part of an interface to a natural language component in a text processing solution. +We are anticipating that a future OVON standard for representing utterances in a dialog system or dialog history will be developed as one example of a derived specification. + +Chapter 1. Specification +1.1 Representation +Dialog event objects will be represented as a JSON [1] object in a string format. The JSON dialog event object is not intended to be a stand-alone document. It is intended to be included in a larger data structure as an object with a re-usable standard structure + +1.2 Dialog Event Object +Each dialog event object has a unique ID, is associated with a single speaker (person or machine), spans a period of time, and contains features representing different inter-related aspects of the event. + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT0.1045" + }, + "features": { + "my-audio-feature": { ... } + "my-text-feature": { ... } + + + "my-semantic-feature": { ... } + "my-custom-feature": { ... } +} } +Figure 1. Example of a dialog event container. +Figure 1 shows an example of a dialog event object. Each dialog event object has a unique id and relates to just one speaker identified by the speaker-id. The event can optionally be linked to a previous event object from the same speaker using the previous-id. The span object represents the time-span of the event. +The features: id, speaker-id, span, and features are mandatory. The previous-id object is optional. The features section will typically contain one or many feature objects but may be empty. Each feature object is identified with a key which can have any arbitrary feature name. As dictionary keys, feature names must be unique within a dialog event. + +1.3 Span +The span object describes a period of time during which the event occurred. It is this span that distinguishes the dialog event from a visual user-interface event which occurs at a point in time. Span objects must contain either an absolute start-time or a relative start-offset but not both. They can also optionally contain either an end-time or end-offset but not both. Absolute times are represented in ISO 8601 [2] format, more specifically the simpler interoperable variant specified in IETF RFC 3339 [3]. Relative durations are represented in ISO 8601 [2] duration format. +The top-level span object may contain a start-offset rather than an absolute start-time. This start-offset specifies the start time as the time elapsed since an absolute reference time is defined outside of the object. For example, the event might be contained in a dialog history object and this offset might be relative to the start-time of the conversation represented by that object. +As will be explained below, span objects can also be included in individual tokens within features of the dialog event. + + +1.4 Feature Objects + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT3.1045" +}, + “features”: { + "my-audio-feature": { + "mime-type": "audio/wav", + "tokens": [ + { + "value-url": "http://localhost/xyz1234.wav" +} ] +}, + "my-text-token-feature": { + "mime-type": "text/plain", + "lang": "en", + "encoding": "UTF-8", + "token-schema": "BertTokenizer.from_pretrained(bert-base-uncased)" + "tokens": [ + {"value": "what"}, + {"value": "is"}, + {"value": "the"}, + {"value": "weather"}, + {"value": "forecast"}, + {"value": "for"}, + {"value": "tomorrow"} +] } +} + + + } +Figure 2. A dialog event containing an audio feature object and a tokenized text feature object. As a minimum, each feature object must contain a mime-type and an array of token objects +named tokens. + +Figure 2 shows a simple example of a tokenized text feature. In this example, the mime-type is plain text and the feature is broken into a sequence of tokenized words. Each word is represented by a token object. +This is just one example of how text can be represented in a feature. The tokens array contains zero or more token objects, and each token object must contain either a value object or a value-url string. The value-url string references an external document that contains the value. The reference is a string in the format of a Universal Resource Locator (URL) as shown in the my-audio-feature in Figure 2. +The mime-type is mandatory and defines the format of the value object or the document referenced by the value-url. The format of mime-type is defined in IETF RFC-2231 [6] +The object lang defines the language of the feature according to IETF BCP 47 language Tag [4]. This value is optional and remains undefined if not included. This value informs how the feature should be interpreted. If the feature contains multilingual content it is recommended that this value remains undefined or is set to the dominant language of the feature. +The encoding object defines the text encoding within the value objects. Values should be one of "ISO-8859-1" (See ISO/IEC 8859-1 [11]) or "UTF-8" (See RFC3629 [12]). This value is optional and remains undefined if not included. +The token-schema object is included to allow reference to standard sets of token values or symbols for certain features with specific mime-types. For example, it could be used to specify which phonetic pronunciation alphabet is present in a pronunciation feature, or refer to a specific standard set of intents for a given domain in a semantic feature. The format of this object is not mandated and is expected to be defined in derived specifications. This value is optional and remains undefined if not included. Figure 2 shows one example of how the token-schema might be used to reference a specific text tokenizer. Figure 4 shows an example of + how token-schema might be used to reference a standard semantic schema. Both examples are illustrative, not normative. + .. +"my-text-token-feature": { +"mime-type": "text/plain", +"tokens": [ + { + "value": "what", + "span": { + "start-offset": "PT0.0210", + "end-offset": "PT0.1457" + } +}, { + "value": "is", + "span": { + "start-offset": "PT0.1460", + "end-offset": "PT0.1976" + } +}, .. { + "value": "tomorrow", + "span": { + "start-offset": "PT2.3082", + "end-offset": "PT3.0784" + } +}, ] +} .. + Figure 3. A feature object containing tokens objects with spans. +Each token object can also optionally contain a span object. This indicates the time span of this particular token. As described in section 1.3 Span, spans can be defined as absolute times or relative offsets. In a token object start-offset and end-offset are relative to the start-time of the parent dialog event (or the implied dialog event start time defined by its start-offset). +Figure 3 shows an example where the tokenized words of the text feature for Figure 2 are each given an individual span. + +1.5 Confidence, Linking, and Stand-Off Annotation +Each feature should represent just one aspect of the dialog event. Each token object within each feature can be linked to another token object in the event using the object links. + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "previous-id": "user-utterance-28", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00", + "end-offset": "PT0.1045" +}, + "features": { + "my-text-feature": { + "mime-type": "text/plain", + "lang": "en", + "encoding": "UTF-8", + "tokens": [ + { + "value": "what is the weather forecast for tomorrow", + "confidence": 0.99 +} ] +}, + "my-semantic-feature": { + "mime-type": "text/x.my-semantic-mime-type", + "token-schema": "standard-intents.org/schemas/weather.xml", + "tokens": [ +{ +"value": { + 9 + + "name": "intent", + "label": "WeatherForecast" + }, + "links": ["$.my-text-feature.tokens[0].value"], + "confidence": 1.0 + }, +{ +"value": { + "name": "date", + "label": "02-14-2023" + }, + "links": [ +"$.my-text-feature.tokens[0].value.substring(34,41)" ], + "confidence": 0.87 + } +] } +} } +Figure 4. A simple example of linking between features with confidence assigned to a feature. +Figure 4 shows an example of a semantic feature which links to the text that it represents. +The optional links object is an array of references to the value (or parts of the value) of other token objects using a JSON Path with some extensions. There is no restriction on how many links a feature can have. +Unlike the text in Figure 2, in this example, the text is not tokenized. The semantic token with the name "intent" is linked to the whole text of the utterance as defined by the JSON path $.my-text-feature.tokens[0].value. The semantic token with the name "date" however is linked to a sub-string of the source text - the word ‘tomorrow’ - via the JSON Path $.my-text-feature.tokens[0].value.substring(34,41). Section 1.6 describes how to use JSON Path references in more detail. +The optional confidence object is a number carrying a measure of the confidence that the information contained in the associated value is ‘correct’. Confidence values are expressed as a probability (real number between 0.0 and 1.0). Derivative specifications may override this range for specialized situations. +Recall that the format of a given value object is defined by the mime-type and will vary between features. Feature layering and cross-referencing are both important parts of the dialog event standard. Together they permit the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. This approach is termed stand-off-annotation. [9] + +1.6 JSON Path and the substring() extension +Individual elements of the linked object are defined as strings containing JSON Paths. JSON Path is a notation for querying and manipulating data stored in JavaScript Object Notation (JSON). +At the time of publication, there is not an agreed standard definition for JSON Path. There are a number of implementations based on JSON Path [7] that are broadly compatible. There is also an IETF task force RFC6901 [8] working on a common standard. +JSON Path expressions in links objects are cross-references from one token object to the value, or part of a value, of another token object. For these references, the root document (designated by ‘$’) is the features object in the current dialog event. There is currently no way to specify a link between features that are not contained in the same dialog event. +Implementations should support one additional non-standard extension to JSON Path - namely the substring() function. This extension is added to allow token objects to reference a specific substring within another token object, for example, a word, phrase, or arbitrary sequence of characters. +The substring function can be invoked at the tail end of a path. The input to this function is the output of the preceding path expression. The substring function takes two ordered comma-separated parameters, the index of the start and end character in the substring. The first character of the string is at the index zero. +For example the JSON Path: $.my-text-feature.tokens[0].value.substring(34,41) will return the 34th to the 41st character of string contained within the object referred to by the JSON Path $.my-text-feature.tokens[0].value. If the referred object is not a string, or the substring is out-of-bounds then the resulting behavior is undefined. + +Token objects can contain content that is specified externally via a value-url string. If this external content is expressed in JSON format then the JSON path should be interpreted as if the value had been specified locally. This is functionally equivalent to reading the content of the document referenced by value-url and placing it in the equivalent local value object before applying the links reference. + +1.7 Alternates + { + "id": "user-utterance-30", + "speaker-id": "b5y09lky5KU5", + "span": { + "start-time": "2022-12-20 15:59:01.246500+00:00" + }, + "features": { + "my-text-feature": { + "mime-type": "text/plain", + "tokens": [ + { + "value": "what is the weather forecast for tomorrow", + "confidence": 0.96 +} ], + "alternates": [ + [ + { + "value": "what is the weather forecast for thursday", + "confidence": 0.73 +} } +] } +} } + 12 + +Figure 5. Defining alternate values for a feature. +Spoken or written language rarely has a single unambiguous interpretation. For this reason, the dialog event object allows for multiple interpretations of a given event. As has already been seen, the tokens object is used to represent the preferred interpretation of a given feature. In addition to this, the optional alternates array can be used to express alternate interpretations of this feature. For example, in Figure 5 it can be seen that the preferred interpretation of the feature named my-text-feature has the value “what is the weather forecast for tomorrow” and a confidence of 0.96. The alternates array contains one additional interpretation of the utterance with the value “what is the weather forecast for thursday” and a confidence of 0.73. +The alternates object is an array of an array. Each object in the outer array represents an alternative interpretation and each inner array represents the token objects for this interpretation. These inner token objects have the same format as array elements in the tokens object. +The outer array can contain zero or more interpretations. An empty alternates array implies that there are no alternate interpretations and is equivalent to the alternates object not being present in the event. + +1.8 Nomenclature +This specification uses ‘kebab-case’ (i.e. hyphenated lowercase) for all nominal property names, for example, start-offset and value-url. User-defined object keys such as the dictionary names used in the features section can use any valid JSON expression. It is recommended but not mandated that any normative key names in derived specifications follow the same kebab-case convention. + + +Chapter 2. Schema +The structure of a JSON dialog event is defined below as a JSON Schema. + { + "$id": "https://openvoicenetwork.org/schema/dialog-event.json", + "$schema": "https://openvoicenetwork.org/schema", + "description": "A representation of a 'language event’ that is to say +any information associated with a phrase, utterance or part of an +utterance.", + "type": "object", + "required": ["id", "speaker-id", "span", "features" ], + "properties": { + "id": { + "type": "string" + }, + "previous-id": { + "type": "string" + }, + "speaker-id": { + "type": "string" + }, + "span": { "$ref": "#/$defs/span" }, + "features": { + "type": "object", + "patternProperties": { + ".*": { "$ref": "#/$defs/features" } + } +} }, + "$defs": { + "features": { + "type": "object", + "required": ["mime-type" , "tokens" ], + "properties": { + "encoding": { + "type": "string", + "description": "The text encoding of the token values" + }, + "mime-type": { + "type": "string", + "description": "The mime-type of the token values" +}, "lang": { + "type": "string", + "description": "The language of the token values" + }, + "token-schema": { + "type": "object", + "description": "A schema restricting the token values" + }, + "tokens" : { + "type": "array", + "items": { + "$ref": "#/$defs/token" + }, + "alternates": { + "type": "array", + "items": { + "type": "array", + "items": { + "type": "object", + "$ref": "#/$defs/token" + } +} } +} } +}, "token": { + "type": "object", + "anyOf": [ + { "required": ["value"] }, + { "required": ["value-url"] } + ], + 15 + + "properties": { + "value": { + "type": ["number","string","object","array","boolean"] + }, + "value-url": { + "$ref": "#/$defs/url" + }, + "confidence": { + "type": "number" + }, + "span": { "$ref": "#/$defs/span" }, + "links" : { + "type": "array", + "items": { + "$ref": "#/$defs/jsonpath" + } +} } +}, "url": { + "type": "string", + "description": "Any valid URL" +}, +"jsonpath": { + "type": "string", + "description": "an expression in JSON Path syntax" +}, "span": { + "anyOf": [ + { "required": ["start-time"] }, + { "required": ["start-offset"] } + ], + "properties": { + "start-time": { + "$ref": "#/$defs/iso-time" + }, + "end-time": { + "$ref": "#/$defs/iso-time" + 16 + + }, + "start-offset": { + "$ref": "#/$defs/iso-duration" + }, + "end-offset": { + "$ref": "#/$defs/iso-duration" +} } + }, + "iso-time": { + "description": "A string in ISO 8601 absolute format." + }, + "iso-duration": { + "description": "A string in ISO 8601 duration format." +} } +} + Chapter 3. References +[1] https://www.ecma-international.org/publications-and-standards/standards/ecma-404/ ECMA-404 The JSON data interchange syntax +[2] https://www.iso.org/iso-8601-date-and-time-format.html ISO 8601 Date and Time Format. +[3] https://datatracker.ietf.org/doc/html/rfc3339 Newman, Chris; Klyne, Graham (July 2002). Date and Time on the Internet: Timestamps. IETF. doi:10.17487/RFC3339. RFC 3339. Archived [4] https://www.rfc-editor.org/rfc/rfc5646.txt RFC 5646. BCP 47. Tags for Identifying Languages. [5] https://www.ietf.org/rfc/rfc4646.txt RFC 4646 Regarding Best Practice for Tags for Identifying Languages +[6] https://datatracker.ietf.org/doc/html/rfc2231 MIME Parameter Value and Encoded Word Extensions: Character Sets, Languages, and Continuations +[7] https://goessner.net/articles/JsonPath/ JSONPath - XPath for JSON. Stefan Goessner. +[8] https://www.rfc-editor.org/rfc/rfc6901 RFC6901. JavaScript Object Notation (JSON) Pointer [9] Pose-Rodriguez, Javier & Lopez, Patrice & Romany, Laurence. (2014). A Generic Formalism for Encoding Standoff annotations in TEI. +[10] https://openvoicenetwork.org/docs/2832-2/ Interoperable Dialog Packet Requirement Specification. +[11] https://www.iso.org/standard/28245.html ISO/IEC 8859-1:1998 Information technology — 8-bit single-byte coded graphic character sets — Part 1: Latin alphabet No. 1 +[12] https://datatracker.ietf.org/doc/html/rfc3629 RFC3629. UTF-8, a transformation format of ISO 10646 + +Chapter 4. Glossary of Terms +Term +Definition +dialog event +A linguistic event in a spoken or written monologue or dialog between two or more speakers. +dialog event object +A JSON object encoding a dialog event as part of an interface to a natural language component in a text or speech processing solution or dialog system. +dialog event feature +A layer of information of a certain type associated with the dialog event. +confidence +A number representing a measure of the confidence that the information contained in the associated value is ‘correct’. +token encoding +The specific encoding used to represent text in a token. +token link +A link from one token in a feature to part or all of another token in a feature used to implement stand-off annotation. +JSON path +An unambiguous reference to part of a JSON object. +language code +A code representing the language (e.g., American English, British English, New Norsk, etc.) +feature mime type +The type of token values in a specific dialog event feature. +dialog event object id +The unique identifier of the dialog event object +token span +Identifies the span of time for an individual token object +dialog event span +Identifies the span of time for the dialog event + +dialog event speaker id +A unique id of the human or machine associated with content of the dialog event +stand-off-annotation +A method of feature layering and cross-referencing that permits the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. +feature token +A representation of part of the information that makes up a feature value. +token object +A JSON object representing a feature token which defines the value of the feature and other associated information such as its span and how it links to other feature tokens. +derived specification +An derivative standard, built upon this specification, which defines a specific way to use a dialog event in a particular context. + +Chapter 5. Decision Log +This section documents some of the key design decisions that were made by the team during the development of this specification. It is informative, not normative. +Issue Topic +Issue and Decision +Nomenclature +Issue: What style of object name should we adopt in the dialog event? +Decision: We adopt ‘kebab-case’ (i.e. hyphenated lower case) for all property names, for example ‘start-offset’ and ‘value-url’. +links +Issue: Should we allow multiple links for a token object? +Decision: Yes so that one feature can reference more than one other feature. +Span start and end times +Issue: Should start times and end times be mandatory or optional? +Decision: Start times are required. End times should be optional because events containing text to be generated via + +Start and end offsets +Time Zones +Features +Speaker Roles +Representation +Representation of Alternates +text-to-speech cannot know when the generated text will actually end in real time. For such generative events the start time will be the time that the utterance is intended to start. It cannot anticipate actual delays in processing. We would anticipate that when the text is processed to generate the audio, the audio will have full span. At this point the end time and the start time may be modified to represent the viewpoint of the downstream process that generated it (i.e. the TTS may add an end time and also change the start time to the point where an event is re-transmitted). +Issue: Should start and end times just be absolute or should we also support relative options such as offsets? +Decision Times can be absolute or relative but not both. The optionality is not complex and supports common use cases. +Issue: Does it matter that ISO can represent different time zones and will it cause a problem if servers are in different time zones? +Decision: No. It shouldn’t matter as long at consumers of messages regularize for their own time zone before using the time. ISO 8601 carries the time zone information in the format. +Issue: What is the cardinality of features object? +Decision: We expect at least one feature but in the spec we will allow for zero or more features. +Issue: Should an event carry speaker roles (e.g. agent, customer, bot, etc.)? +Decision: No. Because the role of a speaker is defined by the consumer of the event not the event itself +Issue: Should we be syntax-neutral or mandate a specific representation? (e.g. YML, JSON and XML) +Decision: We will specify messages as JSON. It is straightforward to translate it at source and destination into other formats. +Should a feature object contain a simple list of lists of tokens or should we separate out alternates into a different object? + Decision: We will treat alternates separately to keep the object simple for systems that do not represent uncertainty in their events. +Feature Cross-Referencing +Issue: Should we use simple array indexing or something more sophisticated such as JSON Path? +Decision: We will use path indexing via JSON Path. Using a path notation means that we do NOT need to insist that values are arrays. They can be any kind of object. This will allow us to reference strings by character indexes OR reference arrays by array indexes OR reference dictionaries by their dictionary names. +External Links +Issue: How do we refer to features in external events? ( This was originally a requirement in the requirements specification.) +Decision: We leave this open for now. See comments on streaming. +Feature Span +How do we attach a time span to a feature? +Decision: We allow the span object to be added to token objects. + OVON Interoperable Event Restrictions + Confidence +Issue: Should we leave the range of confidence values open for derived specifications to decide? +Decision: No. We will mandate a floating point number between 0..1 and leave it up to derived specifications to override this. +Streaming +Issue: Is a stream represented by multiple events in the same stream? +Decision: The use of features and events for streaming needs to be worked through more carefully. It is possible that it will require each feature to have a separate ID and span to allow different spans for different features and to allow continuously evolving interpretations in for example an asynchronous dialog system. + + +Issue: Does streaming apply to both speech and text? Decision: Yes +JPath or JSON Path? +Should we adopt JPath or JSON Path? +Decision: Neither are true standards. JPath appears to be IBM specific. We will adopt JSON Path and extend it to support substring matching. We will also aim to adopt any emerging IETF standard if it is completed. We welcome comments on other standard ways to refer to JSON paths +encoding +Issue: Is it the best approach to follow IETF HTTP in allowing only two encoding types "ISO-8859-1" [ISO-8859-1] or "UTF-8" [RFC3629]. +Decision: Yes for now but we welcome comments. +Token array +Issue: This representation "tokens":[ +{"token": +{ +} }, +{"token": +{ +} } +"value": "what", "span": { +---etc } +"value": "what", "span": { +---etc } +] +is different than the unnamed array elements. This may be valid with JSON indexing. But it permits/forces a JSONtoXML conversion to decide what to "name" the subelements. Decision: to be discussed +Value of “tokens” +Issue: In 1.5 Confidence .... +The value refers to a complete string: +"tokens": [ { +"value": "what is the weather forecast for tomorrow", +"confidence": 0.99 } + ] +But in the earlier example they were single word "tokens". Will +developers need to scan the char-strings to know what they are getting? +Decision: to be discussed +Alternates +Issue: . In the "alternates" section +Figure 5 is missing "]" and if you wanted it to represent +multiple alterates then we should have another alternate array element in the example. +This gets verbose if the "tokens" array is single words as in figure 2. +Decision: to be discussed +Tokens +It looks like according to the Schema the things under "tokens" and "alternates" must be a "token", but the examples don't reflect that. +Decision: to be discussed + From 83a458652bd3ca4058ff4926edb0a27d6c6ad049 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:23:57 -0500 Subject: [PATCH 33/35] Delete Interoperable Dialog Event Object Specification 1.0 --- ...able Dialog Event Object Specification 1.0 | 522 ------------------ 1 file changed, 522 deletions(-) delete mode 100644 specs/Interoperable Dialog Event Object Specification 1.0 diff --git a/specs/Interoperable Dialog Event Object Specification 1.0 b/specs/Interoperable Dialog Event Object Specification 1.0 deleted file mode 100644 index 009b2ca..0000000 --- a/specs/Interoperable Dialog Event Object Specification 1.0 +++ /dev/null @@ -1,522 +0,0 @@ -*********** ---- -layout: default -title: Interoperable Dialog Event Object Specification Version 1.0 -parent: specs - -## 2023.06.09 -## Draft Version 1.0.0 -Interoperable Dialog Event Object Specification Version 1.0 -The Open Voice Network -Architecture Work Group of the Technical Committee -Editor-in-chief: David Attwater Contributors: Emmett Coin -Deborah Dahl Jim Larson - -June 9, 2023 - -TABLE OF CONTENTS -CHAPTER 0. SCOPE AND INTRODUCTION -0.1 Document Scope -0.2 Dialog Events -0.3 A Foundation for Further Specialization - -CHAPTER 1. SPECIFICATION -1.1 Representation -1.2 Dialog Event Object -1.3 Span -1.4 Feature Objects -1.5 Confidence, Linking and Stand-off Annotation -1.6 JSON Path and the substring() extension -1.7 Alternates -1.8 Nomenclature - -CHAPTER 2. SCHEMA -CHAPTER 3. REFERENCES -CHAPTER 4. GLOSSARY OF TERMS -CHAPTER 5. DECISION LOG - - -Chapter 0. Scope and Introduction -0.1 Document Scope -This document specifies the format for Open Voice Network (OVON) interoperable dialog events. The requirements for this specification are given in Interoperable Dialog Packet Requirements [10]. This specification is generic. As described in the requirements there are many different potential uses of a dialog event. - -0.2 Dialog Events -Interoperable conversational systems will need to be able to process linguistic input and generate linguistic output. The components within such systems also need to send and receive such output in a standardized way. -The purpose of a dialog event is to define a generic standardized data structure that can be used in any component of a dialog system to express a ‘language event’, that is to say, any features associated with a phrase, utterance or part of an utterance. Dialog events span a certain time period and are associated with a single speaker. -These events are used to represent user inputs and system outputs of various types, primarily linguistic inputs and outputs such as speech or text, but also potential multimodal inputs and outputs, such as selections on a touchscreen or images presented by a system. Dialog events can be used to express whole utterances, phrases, or other slices of time. -Components of dialog systems may receive an event and add features to it, for example, a natural language interpretation component may receive an event containing a text feature and add a semantic feature to it as an interpretation of the text. -It is anticipated that in the future, events could be joined together to form streams, for example, to represent continuous input or output of a speech-to-text engine. - -0.3 A Foundation for Further Specialization -This specification defines a generic format into which dialog features in different formats can be gathered together into a single event. It also describes how such features can reference each other. -It does not define which formats (mime types) should be present, how features should be named, or which features might be required under what circumstances. This will be left to additional specifications, which we term derived specifications, built upon this one. -For example, dialog events could be used to carry prompt and response information for a dialog system, annotate spoken dialog between two human speakers, or be used as part of an interface to a natural language component in a text processing solution. -We are anticipating that a future OVON standard for representing utterances in a dialog system or dialog history will be developed as one example of a derived specification. - -Chapter 1. Specification -1.1 Representation -Dialog event objects will be represented as a JSON [1] object in a string format. The JSON dialog event object is not intended to be a stand-alone document. It is intended to be included in a larger data structure as an object with a re-usable standard structure - -1.2 Dialog Event Object -Each dialog event object has a unique ID, is associated with a single speaker (person or machine), spans a period of time, and contains features representing different inter-related aspects of the event. - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT0.1045" - }, - "features": { - "my-audio-feature": { ... } - "my-text-feature": { ... } - - - "my-semantic-feature": { ... } - "my-custom-feature": { ... } -} } -Figure 1. Example of a dialog event container. -Figure 1 shows an example of a dialog event object. Each dialog event object has a unique id and relates to just one speaker identified by the speaker-id. The event can optionally be linked to a previous event object from the same speaker using the previous-id. The span object represents the time-span of the event. -The features: id, speaker-id, span, and features are mandatory. The previous-id object is optional. The features section will typically contain one or many feature objects but may be empty. Each feature object is identified with a key which can have any arbitrary feature name. As dictionary keys, feature names must be unique within a dialog event. - -1.3 Span -The span object describes a period of time during which the event occurred. It is this span that distinguishes the dialog event from a visual user-interface event which occurs at a point in time. Span objects must contain either an absolute start-time or a relative start-offset but not both. They can also optionally contain either an end-time or end-offset but not both. Absolute times are represented in ISO 8601 [2] format, more specifically the simpler interoperable variant specified in IETF RFC 3339 [3]. Relative durations are represented in ISO 8601 [2] duration format. -The top-level span object may contain a start-offset rather than an absolute start-time. This start-offset specifies the start time as the time elapsed since an absolute reference time is defined outside of the object. For example, the event might be contained in a dialog history object and this offset might be relative to the start-time of the conversation represented by that object. -As will be explained below, span objects can also be included in individual tokens within features of the dialog event. - - -1.4 Feature Objects - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT3.1045" -}, - “features”: { - "my-audio-feature": { - "mime-type": "audio/wav", - "tokens": [ - { - "value-url": "http://localhost/xyz1234.wav" -} ] -}, - "my-text-token-feature": { - "mime-type": "text/plain", - "lang": "en", - "encoding": "UTF-8", - "token-schema": "BertTokenizer.from_pretrained(bert-base-uncased)" - "tokens": [ - {"value": "what"}, - {"value": "is"}, - {"value": "the"}, - {"value": "weather"}, - {"value": "forecast"}, - {"value": "for"}, - {"value": "tomorrow"} -] } -} - - - } -Figure 2. A dialog event containing an audio feature object and a tokenized text feature object. As a minimum, each feature object must contain a mime-type and an array of token objects -named tokens. - -Figure 2 shows a simple example of a tokenized text feature. In this example, the mime-type is plain text and the feature is broken into a sequence of tokenized words. Each word is represented by a token object. -This is just one example of how text can be represented in a feature. The tokens array contains zero or more token objects, and each token object must contain either a value object or a value-url string. The value-url string references an external document that contains the value. The reference is a string in the format of a Universal Resource Locator (URL) as shown in the my-audio-feature in Figure 2. -The mime-type is mandatory and defines the format of the value object or the document referenced by the value-url. The format of mime-type is defined in IETF RFC-2231 [6] -The object lang defines the language of the feature according to IETF BCP 47 language Tag [4]. This value is optional and remains undefined if not included. This value informs how the feature should be interpreted. If the feature contains multilingual content it is recommended that this value remains undefined or is set to the dominant language of the feature. -The encoding object defines the text encoding within the value objects. Values should be one of "ISO-8859-1" (See ISO/IEC 8859-1 [11]) or "UTF-8" (See RFC3629 [12]). This value is optional and remains undefined if not included. -The token-schema object is included to allow reference to standard sets of token values or symbols for certain features with specific mime-types. For example, it could be used to specify which phonetic pronunciation alphabet is present in a pronunciation feature, or refer to a specific standard set of intents for a given domain in a semantic feature. The format of this object is not mandated and is expected to be defined in derived specifications. This value is optional and remains undefined if not included. Figure 2 shows one example of how the token-schema might be used to reference a specific text tokenizer. Figure 4 shows an example of - how token-schema might be used to reference a standard semantic schema. Both examples are illustrative, not normative. - .. -"my-text-token-feature": { -"mime-type": "text/plain", -"tokens": [ - { - "value": "what", - "span": { - "start-offset": "PT0.0210", - "end-offset": "PT0.1457" - } -}, { - "value": "is", - "span": { - "start-offset": "PT0.1460", - "end-offset": "PT0.1976" - } -}, .. { - "value": "tomorrow", - "span": { - "start-offset": "PT2.3082", - "end-offset": "PT3.0784" - } -}, ] -} .. - Figure 3. A feature object containing tokens objects with spans. -Each token object can also optionally contain a span object. This indicates the time span of this particular token. As described in section 1.3 Span, spans can be defined as absolute times or relative offsets. In a token object start-offset and end-offset are relative to the start-time of the parent dialog event (or the implied dialog event start time defined by its start-offset). -Figure 3 shows an example where the tokenized words of the text feature for Figure 2 are each given an individual span. - -1.5 Confidence, Linking, and Stand-Off Annotation -Each feature should represent just one aspect of the dialog event. Each token object within each feature can be linked to another token object in the event using the object links. - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT0.1045" -}, - "features": { - "my-text-feature": { - "mime-type": "text/plain", - "lang": "en", - "encoding": "UTF-8", - "tokens": [ - { - "value": "what is the weather forecast for tomorrow", - "confidence": 0.99 -} ] -}, - "my-semantic-feature": { - "mime-type": "text/x.my-semantic-mime-type", - "token-schema": "standard-intents.org/schemas/weather.xml", - "tokens": [ -{ -"value": { - 9 - - "name": "intent", - "label": "WeatherForecast" - }, - "links": ["$.my-text-feature.tokens[0].value"], - "confidence": 1.0 - }, -{ -"value": { - "name": "date", - "label": "02-14-2023" - }, - "links": [ -"$.my-text-feature.tokens[0].value.substring(34,41)" ], - "confidence": 0.87 - } -] } -} } -Figure 4. A simple example of linking between features with confidence assigned to a feature. -Figure 4 shows an example of a semantic feature which links to the text that it represents. -The optional links object is an array of references to the value (or parts of the value) of other token objects using a JSON Path with some extensions. There is no restriction on how many links a feature can have. -Unlike the text in Figure 2, in this example, the text is not tokenized. The semantic token with the name "intent" is linked to the whole text of the utterance as defined by the JSON path $.my-text-feature.tokens[0].value. The semantic token with the name "date" however is linked to a sub-string of the source text - the word ‘tomorrow’ - via the JSON Path $.my-text-feature.tokens[0].value.substring(34,41). Section 1.6 describes how to use JSON Path references in more detail. -The optional confidence object is a number carrying a measure of the confidence that the information contained in the associated value is ‘correct’. Confidence values are expressed as a probability (real number between 0.0 and 1.0). Derivative specifications may override this range for specialized situations. -Recall that the format of a given value object is defined by the mime-type and will vary between features. Feature layering and cross-referencing are both important parts of the dialog event standard. Together they permit the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. This approach is termed stand-off-annotation. [9] - -1.6 JSON Path and the substring() extension -Individual elements of the linked object are defined as strings containing JSON Paths. JSON Path is a notation for querying and manipulating data stored in JavaScript Object Notation (JSON). -At the time of publication, there is not an agreed standard definition for JSON Path. There are a number of implementations based on JSON Path [7] that are broadly compatible. There is also an IETF task force RFC6901 [8] working on a common standard. -JSON Path expressions in links objects are cross-references from one token object to the value, or part of a value, of another token object. For these references, the root document (designated by ‘$’) is the features object in the current dialog event. There is currently no way to specify a link between features that are not contained in the same dialog event. -Implementations should support one additional non-standard extension to JSON Path - namely the substring() function. This extension is added to allow token objects to reference a specific substring within another token object, for example, a word, phrase, or arbitrary sequence of characters. -The substring function can be invoked at the tail end of a path. The input to this function is the output of the preceding path expression. The substring function takes two ordered comma-separated parameters, the index of the start and end character in the substring. The first character of the string is at the index zero. -For example the JSON Path: $.my-text-feature.tokens[0].value.substring(34,41) will return the 34th to the 41st character of string contained within the object referred to by the JSON Path $.my-text-feature.tokens[0].value. If the referred object is not a string, or the substring is out-of-bounds then the resulting behavior is undefined. - -Token objects can contain content that is specified externally via a value-url string. If this external content is expressed in JSON format then the JSON path should be interpreted as if the value had been specified locally. This is functionally equivalent to reading the content of the document referenced by value-url and placing it in the equivalent local value object before applying the links reference. - -1.7 Alternates - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00" - }, - "features": { - "my-text-feature": { - "mime-type": "text/plain", - "tokens": [ - { - "value": "what is the weather forecast for tomorrow", - "confidence": 0.96 -} ], - "alternates": [ - [ - { - "value": "what is the weather forecast for thursday", - "confidence": 0.73 -} } -] } -} } - 12 - -Figure 5. Defining alternate values for a feature. -Spoken or written language rarely has a single unambiguous interpretation. For this reason, the dialog event object allows for multiple interpretations of a given event. As has already been seen, the tokens object is used to represent the preferred interpretation of a given feature. In addition to this, the optional alternates array can be used to express alternate interpretations of this feature. For example, in Figure 5 it can be seen that the preferred interpretation of the feature named my-text-feature has the value “what is the weather forecast for tomorrow” and a confidence of 0.96. The alternates array contains one additional interpretation of the utterance with the value “what is the weather forecast for thursday” and a confidence of 0.73. -The alternates object is an array of an array. Each object in the outer array represents an alternative interpretation and each inner array represents the token objects for this interpretation. These inner token objects have the same format as array elements in the tokens object. -The outer array can contain zero or more interpretations. An empty alternates array implies that there are no alternate interpretations and is equivalent to the alternates object not being present in the event. - -1.8 Nomenclature -This specification uses ‘kebab-case’ (i.e. hyphenated lowercase) for all nominal property names, for example, start-offset and value-url. User-defined object keys such as the dictionary names used in the features section can use any valid JSON expression. It is recommended but not mandated that any normative key names in derived specifications follow the same kebab-case convention. - - -Chapter 2. Schema -The structure of a JSON dialog event is defined below as a JSON Schema. - { - "$id": "https://openvoicenetwork.org/schema/dialog-event.json", - "$schema": "https://openvoicenetwork.org/schema", - "description": "A representation of a 'language event’ that is to say -any information associated with a phrase, utterance or part of an -utterance.", - "type": "object", - "required": ["id", "speaker-id", "span", "features" ], - "properties": { - "id": { - "type": "string" - }, - "previous-id": { - "type": "string" - }, - "speaker-id": { - "type": "string" - }, - "span": { "$ref": "#/$defs/span" }, - "features": { - "type": "object", - "patternProperties": { - ".*": { "$ref": "#/$defs/features" } - } -} }, - "$defs": { - "features": { - "type": "object", - "required": ["mime-type" , "tokens" ], - "properties": { - "encoding": { - "type": "string", - "description": "The text encoding of the token values" - }, - "mime-type": { - "type": "string", - "description": "The mime-type of the token values" -}, "lang": { - "type": "string", - "description": "The language of the token values" - }, - "token-schema": { - "type": "object", - "description": "A schema restricting the token values" - }, - "tokens" : { - "type": "array", - "items": { - "$ref": "#/$defs/token" - }, - "alternates": { - "type": "array", - "items": { - "type": "array", - "items": { - "type": "object", - "$ref": "#/$defs/token" - } -} } -} } -}, "token": { - "type": "object", - "anyOf": [ - { "required": ["value"] }, - { "required": ["value-url"] } - ], - 15 - - "properties": { - "value": { - "type": ["number","string","object","array","boolean"] - }, - "value-url": { - "$ref": "#/$defs/url" - }, - "confidence": { - "type": "number" - }, - "span": { "$ref": "#/$defs/span" }, - "links" : { - "type": "array", - "items": { - "$ref": "#/$defs/jsonpath" - } -} } -}, "url": { - "type": "string", - "description": "Any valid URL" -}, -"jsonpath": { - "type": "string", - "description": "an expression in JSON Path syntax" -}, "span": { - "anyOf": [ - { "required": ["start-time"] }, - { "required": ["start-offset"] } - ], - "properties": { - "start-time": { - "$ref": "#/$defs/iso-time" - }, - "end-time": { - "$ref": "#/$defs/iso-time" - 16 - - }, - "start-offset": { - "$ref": "#/$defs/iso-duration" - }, - "end-offset": { - "$ref": "#/$defs/iso-duration" -} } - }, - "iso-time": { - "description": "A string in ISO 8601 absolute format." - }, - "iso-duration": { - "description": "A string in ISO 8601 duration format." -} } -} - Chapter 3. References -[1] https://www.ecma-international.org/publications-and-standards/standards/ecma-404/ ECMA-404 The JSON data interchange syntax -[2] https://www.iso.org/iso-8601-date-and-time-format.html ISO 8601 Date and Time Format. -[3] https://datatracker.ietf.org/doc/html/rfc3339 Newman, Chris; Klyne, Graham (July 2002). Date and Time on the Internet: Timestamps. IETF. doi:10.17487/RFC3339. RFC 3339. Archived [4] https://www.rfc-editor.org/rfc/rfc5646.txt RFC 5646. BCP 47. Tags for Identifying Languages. [5] https://www.ietf.org/rfc/rfc4646.txt RFC 4646 Regarding Best Practice for Tags for Identifying Languages -[6] https://datatracker.ietf.org/doc/html/rfc2231 MIME Parameter Value and Encoded Word Extensions: Character Sets, Languages, and Continuations -[7] https://goessner.net/articles/JsonPath/ JSONPath - XPath for JSON. Stefan Goessner. -[8] https://www.rfc-editor.org/rfc/rfc6901 RFC6901. JavaScript Object Notation (JSON) Pointer [9] Pose-Rodriguez, Javier & Lopez, Patrice & Romany, Laurence. (2014). A Generic Formalism for Encoding Standoff annotations in TEI. -[10] https://openvoicenetwork.org/docs/2832-2/ Interoperable Dialog Packet Requirement Specification. -[11] https://www.iso.org/standard/28245.html ISO/IEC 8859-1:1998 Information technology — 8-bit single-byte coded graphic character sets — Part 1: Latin alphabet No. 1 -[12] https://datatracker.ietf.org/doc/html/rfc3629 RFC3629. UTF-8, a transformation format of ISO 10646 - -Chapter 4. Glossary of Terms -Term -Definition -dialog event -A linguistic event in a spoken or written monologue or dialog between two or more speakers. -dialog event object -A JSON object encoding a dialog event as part of an interface to a natural language component in a text or speech processing solution or dialog system. -dialog event feature -A layer of information of a certain type associated with the dialog event. -confidence -A number representing a measure of the confidence that the information contained in the associated value is ‘correct’. -token encoding -The specific encoding used to represent text in a token. -token link -A link from one token in a feature to part or all of another token in a feature used to implement stand-off annotation. -JSON path -An unambiguous reference to part of a JSON object. -language code -A code representing the language (e.g., American English, British English, New Norsk, etc.) -feature mime type -The type of token values in a specific dialog event feature. -dialog event object id -The unique identifier of the dialog event object -token span -Identifies the span of time for an individual token object -dialog event span -Identifies the span of time for the dialog event - -dialog event speaker id -A unique id of the human or machine associated with content of the dialog event -stand-off-annotation -A method of feature layering and cross-referencing that permits the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. -feature token -A representation of part of the information that makes up a feature value. -token object -A JSON object representing a feature token which defines the value of the feature and other associated information such as its span and how it links to other feature tokens. -derived specification -An derivative standard, built upon this specification, which defines a specific way to use a dialog event in a particular context. - -Chapter 5. Decision Log -This section documents some of the key design decisions that were made by the team during the development of this specification. It is informative, not normative. -Issue Topic -Issue and Decision -Nomenclature -Issue: What style of object name should we adopt in the dialog event? -Decision: We adopt ‘kebab-case’ (i.e. hyphenated lower case) for all property names, for example ‘start-offset’ and ‘value-url’. -links -Issue: Should we allow multiple links for a token object? -Decision: Yes so that one feature can reference more than one other feature. -Span start and end times -Issue: Should start times and end times be mandatory or optional? -Decision: Start times are required. End times should be optional because events containing text to be generated via - -Start and end offsets -Time Zones -Features -Speaker Roles -Representation -Representation of Alternates -text-to-speech cannot know when the generated text will actually end in real time. For such generative events the start time will be the time that the utterance is intended to start. It cannot anticipate actual delays in processing. We would anticipate that when the text is processed to generate the audio, the audio will have full span. At this point the end time and the start time may be modified to represent the viewpoint of the downstream process that generated it (i.e. the TTS may add an end time and also change the start time to the point where an event is re-transmitted). -Issue: Should start and end times just be absolute or should we also support relative options such as offsets? -Decision Times can be absolute or relative but not both. The optionality is not complex and supports common use cases. -Issue: Does it matter that ISO can represent different time zones and will it cause a problem if servers are in different time zones? -Decision: No. It shouldn’t matter as long at consumers of messages regularize for their own time zone before using the time. ISO 8601 carries the time zone information in the format. -Issue: What is the cardinality of features object? -Decision: We expect at least one feature but in the spec we will allow for zero or more features. -Issue: Should an event carry speaker roles (e.g. agent, customer, bot, etc.)? -Decision: No. Because the role of a speaker is defined by the consumer of the event not the event itself -Issue: Should we be syntax-neutral or mandate a specific representation? (e.g. YML, JSON and XML) -Decision: We will specify messages as JSON. It is straightforward to translate it at source and destination into other formats. -Should a feature object contain a simple list of lists of tokens or should we separate out alternates into a different object? - Decision: We will treat alternates separately to keep the object simple for systems that do not represent uncertainty in their events. -Feature Cross-Referencing -Issue: Should we use simple array indexing or something more sophisticated such as JSON Path? -Decision: We will use path indexing via JSON Path. Using a path notation means that we do NOT need to insist that values are arrays. They can be any kind of object. This will allow us to reference strings by character indexes OR reference arrays by array indexes OR reference dictionaries by their dictionary names. -External Links -Issue: How do we refer to features in external events? ( This was originally a requirement in the requirements specification.) -Decision: We leave this open for now. See comments on streaming. -Feature Span -How do we attach a time span to a feature? -Decision: We allow the span object to be added to token objects. - OVON Interoperable Event Restrictions - Confidence -Issue: Should we leave the range of confidence values open for derived specifications to decide? -Decision: No. We will mandate a floating point number between 0..1 and leave it up to derived specifications to override this. -Streaming -Issue: Is a stream represented by multiple events in the same stream? -Decision: The use of features and events for streaming needs to be worked through more carefully. It is possible that it will require each feature to have a separate ID and span to allow different spans for different features and to allow continuously evolving interpretations in for example an asynchronous dialog system. - - -Issue: Does streaming apply to both speech and text? Decision: Yes -JPath or JSON Path? -Should we adopt JPath or JSON Path? -Decision: Neither are true standards. JPath appears to be IBM specific. We will adopt JSON Path and extend it to support substring matching. We will also aim to adopt any emerging IETF standard if it is completed. We welcome comments on other standard ways to refer to JSON paths -encoding -Issue: Is it the best approach to follow IETF HTTP in allowing only two encoding types "ISO-8859-1" [ISO-8859-1] or "UTF-8" [RFC3629]. -Decision: Yes for now but we welcome comments. -Token array -Issue: This representation "tokens":[ -{"token": -{ -} }, -{"token": -{ -} } -"value": "what", "span": { ----etc } -"value": "what", "span": { ----etc } -] -is different than the unnamed array elements. This may be valid with JSON indexing. But it permits/forces a JSONtoXML conversion to decide what to "name" the subelements. Decision: to be discussed -Value of “tokens” -Issue: In 1.5 Confidence .... -The value refers to a complete string: -"tokens": [ { -"value": "what is the weather forecast for tomorrow", -"confidence": 0.99 } - ] -But in the earlier example they were single word "tokens". Will -developers need to scan the char-strings to know what they are getting? -Decision: to be discussed -Alternates -Issue: . In the "alternates" section -Figure 5 is missing "]" and if you wanted it to represent -multiple alterates then we should have another alternate array element in the example. -This gets verbose if the "tokens" array is single words as in figure 2. -Decision: to be discussed -Tokens -It looks like according to the Schema the things under "tokens" and "alternates" must be a "token", but the examples don't reflect that. -Decision: to be discussed - - From 6659c493437d76fb0e1ea58ee461e442d48dc721 Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 15:33:16 -0500 Subject: [PATCH 34/35] Update InteropDialogEventSpecs.md --- specs/InteropDialogEventSpecs.md | 54 +++++++++++++++++--------------- 1 file changed, 28 insertions(+), 26 deletions(-) diff --git a/specs/InteropDialogEventSpecs.md b/specs/InteropDialogEventSpecs.md index eda6d60..679ea26 100644 --- a/specs/InteropDialogEventSpecs.md +++ b/specs/InteropDialogEventSpecs.md @@ -6,34 +6,36 @@ parent: specs ## 2023.06.09 ## Draft Version 1.0.0 -Interoperable Dialog Event Object Specification Version 1.0 -The Open Voice Network -Architecture Work Group of the Technical Committee -Editor-in-chief: David Attwater Contributors: Emmett Coin -Deborah Dahl Jim Larson -June 9, 2023 -TABLE OF CONTENTS -CHAPTER 0. SCOPE AND INTRODUCTION -0.1 Document Scope -0.2 Dialog Events -0.3 A Foundation for Further Specialization - -CHAPTER 1. SPECIFICATION -1.1 Representation -1.2 Dialog Event Object -1.3 Span -1.4 Feature Objects -1.5 Confidence, Linking and Stand-off Annotation -1.6 JSON Path and the substring() extension -1.7 Alternates -1.8 Nomenclature - -CHAPTER 2. SCHEMA -CHAPTER 3. REFERENCES -CHAPTER 4. GLOSSARY OF TERMS -CHAPTER 5. DECISION LOG +**Interoperable Dialog Event Object Specification Version 1.0** +**The Open Voice Network** +**Architecture Work Group of the Technical Committee** +**Editor-in-chief: David Attwater** +**Contributors: Emmett Coin, Deborah Dahl, Jim Larson** + +**June 9, 2023** + +# TABLE OF CONTENTS +## CHAPTER 0. SCOPE AND INTRODUCTION +### 0.1 Document Scope +### 0.2 Dialog Events +### 0.3 A Foundation for Further Specialization + +## CHAPTER 1. SPECIFICATION +### 1.1 Representation +### 1.2 Dialog Event Object +### 1.3 Span +### 1.4 Feature Objects +### 1.5 Confidence, Linking and Stand-off Annotation +### 1.6 JSON Path and the substring() extension +### 1.7 Alternates +### 1.8 Nomenclature + +## CHAPTER 2. SCHEMA +## CHAPTER 3. REFERENCES +## CHAPTER 4. GLOSSARY OF TERMS +## CHAPTER 5. DECISION LOG Chapter 0. Scope and Introduction From 6b7a04f1a257ffb21e5ea4ad9dc563ce4826c16c Mon Sep 17 00:00:00 2001 From: NSouthernLF <101130471+NSouthernLF@users.noreply.github.com> Date: Thu, 15 Jun 2023 22:54:02 -0500 Subject: [PATCH 35/35] Update InteropDialogEventSpecs.md --- specs/InteropDialogEventSpecs.md | 659 ++++++++----------------------- 1 file changed, 171 insertions(+), 488 deletions(-) diff --git a/specs/InteropDialogEventSpecs.md b/specs/InteropDialogEventSpecs.md index 679ea26..8a3726c 100644 --- a/specs/InteropDialogEventSpecs.md +++ b/specs/InteropDialogEventSpecs.md @@ -1,14 +1,13 @@ *********** --- layout: default -title: Interoperable Dialog Event Object Specification Version 1.0 +Interoperable Dialog Event Object Specification Version 1.0 parent: specs ## 2023.06.09 ## Draft Version 1.0.0 - -**Interoperable Dialog Event Object Specification Version 1.0** +# Interoperable Dialog Event Object Specification Version 1.0 **The Open Voice Network** **Architecture Work Group of the Technical Committee** **Editor-in-chief: David Attwater** @@ -16,11 +15,13 @@ parent: specs **June 9, 2023** -# TABLE OF CONTENTS +## TABLE OF CONTENTS + + ## CHAPTER 0. SCOPE AND INTRODUCTION -### 0.1 Document Scope +### 0.1 Document Scope ### 0.2 Dialog Events -### 0.3 A Foundation for Further Specialization +### 0.3 A Foundation for Further Specialization ## CHAPTER 1. SPECIFICATION ### 1.1 Representation @@ -38,486 +39,168 @@ parent: specs ## CHAPTER 5. DECISION LOG -Chapter 0. Scope and Introduction -0.1 Document Scope -This document specifies the format for Open Voice Network (OVON) interoperable dialog events. The requirements for this specification are given in Interoperable Dialog Packet Requirements [10]. This specification is generic. As described in the requirements there are many different potential uses of a dialog event. - -0.2 Dialog Events -Interoperable conversational systems will need to be able to process linguistic input and generate linguistic output. The components within such systems also need to send and receive such output in a standardized way. -The purpose of a dialog event is to define a generic standardized data structure that can be used in any component of a dialog system to express a ‘language event’, that is to say, any features associated with a phrase, utterance or part of an utterance. Dialog events span a certain time period and are associated with a single speaker. -These events are used to represent user inputs and system outputs of various types, primarily linguistic inputs and outputs such as speech or text, but also potential multimodal inputs and outputs, such as selections on a touchscreen or images presented by a system. Dialog events can be used to express whole utterances, phrases, or other slices of time. -Components of dialog systems may receive an event and add features to it, for example, a natural language interpretation component may receive an event containing a text feature and add a semantic feature to it as an interpretation of the text. -It is anticipated that in the future, events could be joined together to form streams, for example, to represent continuous input or output of a speech-to-text engine. - -0.3 A Foundation for Further Specialization -This specification defines a generic format into which dialog features in different formats can be gathered together into a single event. It also describes how such features can reference each other. -It does not define which formats (mime types) should be present, how features should be named, or which features might be required under what circumstances. This will be left to additional specifications, which we term derived specifications, built upon this one. -For example, dialog events could be used to carry prompt and response information for a dialog system, annotate spoken dialog between two human speakers, or be used as part of an interface to a natural language component in a text processing solution. -We are anticipating that a future OVON standard for representing utterances in a dialog system or dialog history will be developed as one example of a derived specification. - -Chapter 1. Specification -1.1 Representation -Dialog event objects will be represented as a JSON [1] object in a string format. The JSON dialog event object is not intended to be a stand-alone document. It is intended to be included in a larger data structure as an object with a re-usable standard structure - -1.2 Dialog Event Object -Each dialog event object has a unique ID, is associated with a single speaker (person or machine), spans a period of time, and contains features representing different inter-related aspects of the event. - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT0.1045" - }, - "features": { - "my-audio-feature": { ... } - "my-text-feature": { ... } - - - "my-semantic-feature": { ... } - "my-custom-feature": { ... } -} } -Figure 1. Example of a dialog event container. -Figure 1 shows an example of a dialog event object. Each dialog event object has a unique id and relates to just one speaker identified by the speaker-id. The event can optionally be linked to a previous event object from the same speaker using the previous-id. The span object represents the time-span of the event. -The features: id, speaker-id, span, and features are mandatory. The previous-id object is optional. The features section will typically contain one or many feature objects but may be empty. Each feature object is identified with a key which can have any arbitrary feature name. As dictionary keys, feature names must be unique within a dialog event. - -1.3 Span -The span object describes a period of time during which the event occurred. It is this span that distinguishes the dialog event from a visual user-interface event which occurs at a point in time. Span objects must contain either an absolute start-time or a relative start-offset but not both. They can also optionally contain either an end-time or end-offset but not both. Absolute times are represented in ISO 8601 [2] format, more specifically the simpler interoperable variant specified in IETF RFC 3339 [3]. Relative durations are represented in ISO 8601 [2] duration format. -The top-level span object may contain a start-offset rather than an absolute start-time. This start-offset specifies the start time as the time elapsed since an absolute reference time is defined outside of the object. For example, the event might be contained in a dialog history object and this offset might be relative to the start-time of the conversation represented by that object. -As will be explained below, span objects can also be included in individual tokens within features of the dialog event. - - -1.4 Feature Objects - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT3.1045" -}, - “features”: { - "my-audio-feature": { - "mime-type": "audio/wav", - "tokens": [ - { - "value-url": "http://localhost/xyz1234.wav" -} ] -}, - "my-text-token-feature": { - "mime-type": "text/plain", - "lang": "en", - "encoding": "UTF-8", - "token-schema": "BertTokenizer.from_pretrained(bert-base-uncased)" - "tokens": [ - {"value": "what"}, - {"value": "is"}, - {"value": "the"}, - {"value": "weather"}, - {"value": "forecast"}, - {"value": "for"}, - {"value": "tomorrow"} -] } -} - - - } -Figure 2. A dialog event containing an audio feature object and a tokenized text feature object. As a minimum, each feature object must contain a mime-type and an array of token objects -named tokens. - -Figure 2 shows a simple example of a tokenized text feature. In this example, the mime-type is plain text and the feature is broken into a sequence of tokenized words. Each word is represented by a token object. -This is just one example of how text can be represented in a feature. The tokens array contains zero or more token objects, and each token object must contain either a value object or a value-url string. The value-url string references an external document that contains the value. The reference is a string in the format of a Universal Resource Locator (URL) as shown in the my-audio-feature in Figure 2. -The mime-type is mandatory and defines the format of the value object or the document referenced by the value-url. The format of mime-type is defined in IETF RFC-2231 [6] -The object lang defines the language of the feature according to IETF BCP 47 language Tag [4]. This value is optional and remains undefined if not included. This value informs how the feature should be interpreted. If the feature contains multilingual content it is recommended that this value remains undefined or is set to the dominant language of the feature. -The encoding object defines the text encoding within the value objects. Values should be one of "ISO-8859-1" (See ISO/IEC 8859-1 [11]) or "UTF-8" (See RFC3629 [12]). This value is optional and remains undefined if not included. -The token-schema object is included to allow reference to standard sets of token values or symbols for certain features with specific mime-types. For example, it could be used to specify which phonetic pronunciation alphabet is present in a pronunciation feature, or refer to a specific standard set of intents for a given domain in a semantic feature. The format of this object is not mandated and is expected to be defined in derived specifications. This value is optional and remains undefined if not included. Figure 2 shows one example of how the token-schema might be used to reference a specific text tokenizer. Figure 4 shows an example of - how token-schema might be used to reference a standard semantic schema. Both examples are illustrative, not normative. - .. -"my-text-token-feature": { -"mime-type": "text/plain", -"tokens": [ - { - "value": "what", - "span": { - "start-offset": "PT0.0210", - "end-offset": "PT0.1457" - } -}, { - "value": "is", - "span": { - "start-offset": "PT0.1460", - "end-offset": "PT0.1976" - } -}, .. { - "value": "tomorrow", - "span": { - "start-offset": "PT2.3082", - "end-offset": "PT3.0784" - } -}, ] -} .. - Figure 3. A feature object containing tokens objects with spans. -Each token object can also optionally contain a span object. This indicates the time span of this particular token. As described in section 1.3 Span, spans can be defined as absolute times or relative offsets. In a token object start-offset and end-offset are relative to the start-time of the parent dialog event (or the implied dialog event start time defined by its start-offset). -Figure 3 shows an example where the tokenized words of the text feature for Figure 2 are each given an individual span. - -1.5 Confidence, Linking, and Stand-Off Annotation -Each feature should represent just one aspect of the dialog event. Each token object within each feature can be linked to another token object in the event using the object links. - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "previous-id": "user-utterance-28", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00", - "end-offset": "PT0.1045" -}, - "features": { - "my-text-feature": { - "mime-type": "text/plain", - "lang": "en", - "encoding": "UTF-8", - "tokens": [ - { - "value": "what is the weather forecast for tomorrow", - "confidence": 0.99 -} ] -}, - "my-semantic-feature": { - "mime-type": "text/x.my-semantic-mime-type", - "token-schema": "standard-intents.org/schemas/weather.xml", - "tokens": [ -{ -"value": { - 9 - - "name": "intent", - "label": "WeatherForecast" - }, - "links": ["$.my-text-feature.tokens[0].value"], - "confidence": 1.0 - }, -{ -"value": { - "name": "date", - "label": "02-14-2023" - }, - "links": [ -"$.my-text-feature.tokens[0].value.substring(34,41)" ], - "confidence": 0.87 - } -] } -} } -Figure 4. A simple example of linking between features with confidence assigned to a feature. -Figure 4 shows an example of a semantic feature which links to the text that it represents. -The optional links object is an array of references to the value (or parts of the value) of other token objects using a JSON Path with some extensions. There is no restriction on how many links a feature can have. -Unlike the text in Figure 2, in this example, the text is not tokenized. The semantic token with the name "intent" is linked to the whole text of the utterance as defined by the JSON path $.my-text-feature.tokens[0].value. The semantic token with the name "date" however is linked to a sub-string of the source text - the word ‘tomorrow’ - via the JSON Path $.my-text-feature.tokens[0].value.substring(34,41). Section 1.6 describes how to use JSON Path references in more detail. -The optional confidence object is a number carrying a measure of the confidence that the information contained in the associated value is ‘correct’. Confidence values are expressed as a probability (real number between 0.0 and 1.0). Derivative specifications may override this range for specialized situations. -Recall that the format of a given value object is defined by the mime-type and will vary between features. Feature layering and cross-referencing are both important parts of the dialog event standard. Together they permit the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. This approach is termed stand-off-annotation. [9] - -1.6 JSON Path and the substring() extension -Individual elements of the linked object are defined as strings containing JSON Paths. JSON Path is a notation for querying and manipulating data stored in JavaScript Object Notation (JSON). -At the time of publication, there is not an agreed standard definition for JSON Path. There are a number of implementations based on JSON Path [7] that are broadly compatible. There is also an IETF task force RFC6901 [8] working on a common standard. -JSON Path expressions in links objects are cross-references from one token object to the value, or part of a value, of another token object. For these references, the root document (designated by ‘$’) is the features object in the current dialog event. There is currently no way to specify a link between features that are not contained in the same dialog event. -Implementations should support one additional non-standard extension to JSON Path - namely the substring() function. This extension is added to allow token objects to reference a specific substring within another token object, for example, a word, phrase, or arbitrary sequence of characters. -The substring function can be invoked at the tail end of a path. The input to this function is the output of the preceding path expression. The substring function takes two ordered comma-separated parameters, the index of the start and end character in the substring. The first character of the string is at the index zero. -For example the JSON Path: $.my-text-feature.tokens[0].value.substring(34,41) will return the 34th to the 41st character of string contained within the object referred to by the JSON Path $.my-text-feature.tokens[0].value. If the referred object is not a string, or the substring is out-of-bounds then the resulting behavior is undefined. - -Token objects can contain content that is specified externally via a value-url string. If this external content is expressed in JSON format then the JSON path should be interpreted as if the value had been specified locally. This is functionally equivalent to reading the content of the document referenced by value-url and placing it in the equivalent local value object before applying the links reference. - -1.7 Alternates - { - "id": "user-utterance-30", - "speaker-id": "b5y09lky5KU5", - "span": { - "start-time": "2022-12-20 15:59:01.246500+00:00" - }, - "features": { - "my-text-feature": { - "mime-type": "text/plain", - "tokens": [ - { - "value": "what is the weather forecast for tomorrow", - "confidence": 0.96 -} ], - "alternates": [ - [ - { - "value": "what is the weather forecast for thursday", - "confidence": 0.73 -} } -] } -} } - 12 - -Figure 5. Defining alternate values for a feature. -Spoken or written language rarely has a single unambiguous interpretation. For this reason, the dialog event object allows for multiple interpretations of a given event. As has already been seen, the tokens object is used to represent the preferred interpretation of a given feature. In addition to this, the optional alternates array can be used to express alternate interpretations of this feature. For example, in Figure 5 it can be seen that the preferred interpretation of the feature named my-text-feature has the value “what is the weather forecast for tomorrow” and a confidence of 0.96. The alternates array contains one additional interpretation of the utterance with the value “what is the weather forecast for thursday” and a confidence of 0.73. -The alternates object is an array of an array. Each object in the outer array represents an alternative interpretation and each inner array represents the token objects for this interpretation. These inner token objects have the same format as array elements in the tokens object. -The outer array can contain zero or more interpretations. An empty alternates array implies that there are no alternate interpretations and is equivalent to the alternates object not being present in the event. - -1.8 Nomenclature -This specification uses ‘kebab-case’ (i.e. hyphenated lowercase) for all nominal property names, for example, start-offset and value-url. User-defined object keys such as the dictionary names used in the features section can use any valid JSON expression. It is recommended but not mandated that any normative key names in derived specifications follow the same kebab-case convention. - - -Chapter 2. Schema -The structure of a JSON dialog event is defined below as a JSON Schema. - { - "$id": "https://openvoicenetwork.org/schema/dialog-event.json", - "$schema": "https://openvoicenetwork.org/schema", - "description": "A representation of a 'language event’ that is to say -any information associated with a phrase, utterance or part of an -utterance.", - "type": "object", - "required": ["id", "speaker-id", "span", "features" ], - "properties": { - "id": { - "type": "string" - }, - "previous-id": { - "type": "string" - }, - "speaker-id": { - "type": "string" - }, - "span": { "$ref": "#/$defs/span" }, - "features": { - "type": "object", - "patternProperties": { - ".*": { "$ref": "#/$defs/features" } - } -} }, - "$defs": { - "features": { - "type": "object", - "required": ["mime-type" , "tokens" ], - "properties": { - "encoding": { - "type": "string", - "description": "The text encoding of the token values" - }, - "mime-type": { - "type": "string", - "description": "The mime-type of the token values" -}, "lang": { - "type": "string", - "description": "The language of the token values" - }, - "token-schema": { - "type": "object", - "description": "A schema restricting the token values" - }, - "tokens" : { - "type": "array", - "items": { - "$ref": "#/$defs/token" - }, - "alternates": { - "type": "array", - "items": { - "type": "array", - "items": { - "type": "object", - "$ref": "#/$defs/token" - } -} } -} } -}, "token": { - "type": "object", - "anyOf": [ - { "required": ["value"] }, - { "required": ["value-url"] } - ], - 15 - - "properties": { - "value": { - "type": ["number","string","object","array","boolean"] - }, - "value-url": { - "$ref": "#/$defs/url" - }, - "confidence": { - "type": "number" - }, - "span": { "$ref": "#/$defs/span" }, - "links" : { - "type": "array", - "items": { - "$ref": "#/$defs/jsonpath" - } -} } -}, "url": { - "type": "string", - "description": "Any valid URL" -}, -"jsonpath": { - "type": "string", - "description": "an expression in JSON Path syntax" -}, "span": { - "anyOf": [ - { "required": ["start-time"] }, - { "required": ["start-offset"] } - ], - "properties": { - "start-time": { - "$ref": "#/$defs/iso-time" - }, - "end-time": { - "$ref": "#/$defs/iso-time" - 16 - - }, - "start-offset": { - "$ref": "#/$defs/iso-duration" - }, - "end-offset": { - "$ref": "#/$defs/iso-duration" -} } - }, - "iso-time": { - "description": "A string in ISO 8601 absolute format." - }, - "iso-duration": { - "description": "A string in ISO 8601 duration format." -} } -} - Chapter 3. References -[1] https://www.ecma-international.org/publications-and-standards/standards/ecma-404/ ECMA-404 The JSON data interchange syntax -[2] https://www.iso.org/iso-8601-date-and-time-format.html ISO 8601 Date and Time Format. -[3] https://datatracker.ietf.org/doc/html/rfc3339 Newman, Chris; Klyne, Graham (July 2002). Date and Time on the Internet: Timestamps. IETF. doi:10.17487/RFC3339. RFC 3339. Archived [4] https://www.rfc-editor.org/rfc/rfc5646.txt RFC 5646. BCP 47. Tags for Identifying Languages. [5] https://www.ietf.org/rfc/rfc4646.txt RFC 4646 Regarding Best Practice for Tags for Identifying Languages -[6] https://datatracker.ietf.org/doc/html/rfc2231 MIME Parameter Value and Encoded Word Extensions: Character Sets, Languages, and Continuations -[7] https://goessner.net/articles/JsonPath/ JSONPath - XPath for JSON. Stefan Goessner. -[8] https://www.rfc-editor.org/rfc/rfc6901 RFC6901. JavaScript Object Notation (JSON) Pointer [9] Pose-Rodriguez, Javier & Lopez, Patrice & Romany, Laurence. (2014). A Generic Formalism for Encoding Standoff annotations in TEI. -[10] https://openvoicenetwork.org/docs/2832-2/ Interoperable Dialog Packet Requirement Specification. -[11] https://www.iso.org/standard/28245.html ISO/IEC 8859-1:1998 Information technology — 8-bit single-byte coded graphic character sets — Part 1: Latin alphabet No. 1 -[12] https://datatracker.ietf.org/doc/html/rfc3629 RFC3629. UTF-8, a transformation format of ISO 10646 - -Chapter 4. Glossary of Terms -Term -Definition -dialog event -A linguistic event in a spoken or written monologue or dialog between two or more speakers. -dialog event object -A JSON object encoding a dialog event as part of an interface to a natural language component in a text or speech processing solution or dialog system. -dialog event feature -A layer of information of a certain type associated with the dialog event. -confidence -A number representing a measure of the confidence that the information contained in the associated value is ‘correct’. -token encoding -The specific encoding used to represent text in a token. -token link -A link from one token in a feature to part or all of another token in a feature used to implement stand-off annotation. -JSON path -An unambiguous reference to part of a JSON object. -language code -A code representing the language (e.g., American English, British English, New Norsk, etc.) -feature mime type -The type of token values in a specific dialog event feature. -dialog event object id -The unique identifier of the dialog event object -token span -Identifies the span of time for an individual token object -dialog event span -Identifies the span of time for the dialog event - -dialog event speaker id -A unique id of the human or machine associated with content of the dialog event -stand-off-annotation -A method of feature layering and cross-referencing that permits the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. -feature token -A representation of part of the information that makes up a feature value. -token object -A JSON object representing a feature token which defines the value of the feature and other associated information such as its span and how it links to other feature tokens. -derived specification -An derivative standard, built upon this specification, which defines a specific way to use a dialog event in a particular context. - -Chapter 5. Decision Log -This section documents some of the key design decisions that were made by the team during the development of this specification. It is informative, not normative. -Issue Topic -Issue and Decision -Nomenclature -Issue: What style of object name should we adopt in the dialog event? -Decision: We adopt ‘kebab-case’ (i.e. hyphenated lower case) for all property names, for example ‘start-offset’ and ‘value-url’. -links -Issue: Should we allow multiple links for a token object? -Decision: Yes so that one feature can reference more than one other feature. -Span start and end times -Issue: Should start times and end times be mandatory or optional? -Decision: Start times are required. End times should be optional because events containing text to be generated via - -Start and end offsets -Time Zones -Features -Speaker Roles -Representation -Representation of Alternates -text-to-speech cannot know when the generated text will actually end in real time. For such generative events the start time will be the time that the utterance is intended to start. It cannot anticipate actual delays in processing. We would anticipate that when the text is processed to generate the audio, the audio will have full span. At this point the end time and the start time may be modified to represent the viewpoint of the downstream process that generated it (i.e. the TTS may add an end time and also change the start time to the point where an event is re-transmitted). -Issue: Should start and end times just be absolute or should we also support relative options such as offsets? -Decision Times can be absolute or relative but not both. The optionality is not complex and supports common use cases. -Issue: Does it matter that ISO can represent different time zones and will it cause a problem if servers are in different time zones? -Decision: No. It shouldn’t matter as long at consumers of messages regularize for their own time zone before using the time. ISO 8601 carries the time zone information in the format. -Issue: What is the cardinality of features object? -Decision: We expect at least one feature but in the spec we will allow for zero or more features. -Issue: Should an event carry speaker roles (e.g. agent, customer, bot, etc.)? -Decision: No. Because the role of a speaker is defined by the consumer of the event not the event itself -Issue: Should we be syntax-neutral or mandate a specific representation? (e.g. YML, JSON and XML) -Decision: We will specify messages as JSON. It is straightforward to translate it at source and destination into other formats. -Should a feature object contain a simple list of lists of tokens or should we separate out alternates into a different object? - Decision: We will treat alternates separately to keep the object simple for systems that do not represent uncertainty in their events. -Feature Cross-Referencing -Issue: Should we use simple array indexing or something more sophisticated such as JSON Path? -Decision: We will use path indexing via JSON Path. Using a path notation means that we do NOT need to insist that values are arrays. They can be any kind of object. This will allow us to reference strings by character indexes OR reference arrays by array indexes OR reference dictionaries by their dictionary names. -External Links -Issue: How do we refer to features in external events? ( This was originally a requirement in the requirements specification.) -Decision: We leave this open for now. See comments on streaming. -Feature Span -How do we attach a time span to a feature? -Decision: We allow the span object to be added to token objects. - OVON Interoperable Event Restrictions - Confidence -Issue: Should we leave the range of confidence values open for derived specifications to decide? -Decision: No. We will mandate a floating point number between 0..1 and leave it up to derived specifications to override this. -Streaming -Issue: Is a stream represented by multiple events in the same stream? -Decision: The use of features and events for streaming needs to be worked through more carefully. It is possible that it will require each feature to have a separate ID and span to allow different spans for different features and to allow continuously evolving interpretations in for example an asynchronous dialog system. - - -Issue: Does streaming apply to both speech and text? Decision: Yes -JPath or JSON Path? -Should we adopt JPath or JSON Path? -Decision: Neither are true standards. JPath appears to be IBM specific. We will adopt JSON Path and extend it to support substring matching. We will also aim to adopt any emerging IETF standard if it is completed. We welcome comments on other standard ways to refer to JSON paths -encoding -Issue: Is it the best approach to follow IETF HTTP in allowing only two encoding types "ISO-8859-1" [ISO-8859-1] or "UTF-8" [RFC3629]. -Decision: Yes for now but we welcome comments. -Token array -Issue: This representation "tokens":[ -{"token": -{ -} }, -{"token": -{ -} } -"value": "what", "span": { ----etc } -"value": "what", "span": { ----etc } -] -is different than the unnamed array elements. This may be valid with JSON indexing. But it permits/forces a JSONtoXML conversion to decide what to "name" the subelements. Decision: to be discussed -Value of “tokens” -Issue: In 1.5 Confidence .... -The value refers to a complete string: -"tokens": [ { -"value": "what is the weather forecast for tomorrow", -"confidence": 0.99 } - ] -But in the earlier example they were single word "tokens". Will -developers need to scan the char-strings to know what they are getting? -Decision: to be discussed -Alternates -Issue: . In the "alternates" section -Figure 5 is missing "]" and if you wanted it to represent -multiple alterates then we should have another alternate array element in the example. -This gets verbose if the "tokens" array is single words as in figure 2. -Decision: to be discussed -Tokens -It looks like according to the Schema the things under "tokens" and "alternates" must be a "token", but the examples don't reflect that. -Decision: to be discussed +## Chapter 0. Scope and Introduction + +### 0.1 Document Scope + +##### This document specifies the format for Open Voice Network (OVON) interoperable dialog events. The requirements for this specification are given in [Interoperable Dialog Packet Requirements [10]](https://openvoicenetwork.org/docs/2832-2/). This specification is generic. As described in the requirements there are many different potential uses of a dialog event. ##### + +### 0.2 Dialog Events + +##### Interoperable conversational systems will need to be able to process linguistic input and generate linguistic output. The components within such systems also need to send and receive such output in a standardized way. ##### + +##### The purpose of a dialog event is to define a generic standardized data structure that can be used in any component of a dialog system to express a ‘language event’, that is to say, any features associated with a phrase, utterance or part of an utterance. Dialog events span a certain time period and are associated with a single speaker. ##### + +##### These events are used to represent user inputs and system outputs of various types, primarily linguistic inputs and outputs such as speech or text, but also potential multimodal inputs and outputs, such as selections on a touchscreen or images presented by a system. Dialog events can be used to express whole utterances, phrases, or other slices of time. ##### + +##### Components of dialog systems may receive an event and add features to it, for example, a natural language interpretation component may receive an event containing a text feature and add a semantic feature to it as an interpretation of the text. ##### + +##### It is anticipated that in the future, events could be joined together to form streams, for example, to represent continuous input or output of a speech-to-text engine. ##### + +### 0.3 A Foundation for Further Specialization + +##### This specification defines a generic format into which dialog features in different formats can be gathered together into a single event. It also describes how such features can reference each other. ##### + +##### It does not define which formats (mime types) should be present, how features should be named, or which features might be required under what circumstances. This will be left to additional specifications, which we term _derived specifications_, built upon this one. ##### + +##### For example, dialog events could be used to carry prompt and response information for a dialog system, annotate spoken dialog between two human speakers, or be used as part of an interface to a natural language component in a text processing solution. ##### + +##### We are anticipating that a future OVON standard for representing utterances in a dialog system or dialog history will be developed as one example of a derived specification. ##### + +## Chapter 1. Specification +### 1.1 Representation +##### Dialog event objects will be represented as a JSON [1] object in a string format. The JSON dialog event object is not intended to be a stand-alone document. It is intended to be included in a larger data structure as an object with a re-usable standard structure + +### 1.2 Dialog Event Object + +##### Each dialog event object has a unique ID, is associated with a single speaker (person or machine), spans a period of time, and contains features representing different inter-related aspects of the event. ##### + +![diagram1](https://github.com/NSouthernLF/docs/assets/101130471/c55c66e2-79e8-4f16-bc9c-f0e81728821e) + +##### _Figure 1. Example of a dialog event container._ ##### + +##### Figure 1 shows an example of a dialog event object. Each dialog event object has a unique id and relates to just one speaker identified by the speaker-id. The event can optionally be linked to a previous event object from the same speaker using the previous-id. The span object represents the time-span of the event. ##### + +##### The features: id, speaker-id, span, and features are mandatory. The previous-id object is optional. The features section will typically contain one or many feature objects but may be empty. Each feature object is identified with a key which can have any arbitrary feature name. As dictionary keys, feature names must be unique within a dialog event. ##### + + +### 1.3 Span +##### The _span_ object describes a period of time during which the event occurred. It is this span that distinguishes the dialog event from a visual user-interface event which occurs at a point in time. Span objects must contain either an absolute _start-time_ or a relative _start-offset_ but not both. They can also optionally contain either an _end-time_ or _end-offset_ but not both. Absolute times are represented in [ISO 8601 [2]](https://www.iso.org/iso-8601-date-and-time-format.html) format, more specifically the simpler interoperable variant specified in [IETF RFC 3339 [3]](https://datatracker.ietf.org/doc/html/rfc3339). Relative durations are represented in [ISO 8601 [2]](https://www.iso.org/iso-8601-date-and-time-format.html) duration format. ##### + +##### The top-level span object may contain a _start-offset_ rather than an absolute _start-time_. This _start-offset_ specifies the start time as the time elapsed since an absolute reference time is defined outside of the object. For example, the event might be contained in a dialog history object and this offset might be relative to the start-time of the conversation represented by that object. ##### + +##### As will be explained below, span objects can also be included in individual tokens within features of the dialog event. ##### + + +### 1.4 Feature Objects +![diagram2](https://github.com/NSouthernLF/docs/assets/101130471/b036994b-44c1-46b7-a8d5-35337c69bbb8) +##### _Figure 2. A dialog event containing an audio feature object and a tokenized text feature object._ ##### + +##### As a minimum, each feature object must contain a _mime-type_ and an array of token objects named _tokens_ ##### + + +##### Figure 2 shows a simple example of a tokenized text feature. In this example, the _mime-type_ is plain text and the feature is broken into a sequence of tokenized words. Each word is represented by a _token_ object. ##### + +##### This is just one example of how text can be represented in a feature. The _tokens_ array contains zero or more token objects, and each token object must contain either a value object or a value-url string. The value-url string references an external document that contains the value. The reference is a string in the format of a Universal Resource Locator (URL) as shown in the my-audio-feature in Figure 2. ##### + +##### The _mime-type_ is mandatory and defines the format of the _value_ object or the document referenced by the _value-url_. The format of _mime-type_ is defined in IETF RFC-2231 [6] ##### + +##### The object _lang_ defines the language of the feature according to [IETF BCP 47 language Tag [4]](https://www.rfc-editor.org/rfc/rfc5646.txt). This value is optional and remains undefined if not included. This value informs how the feature should be interpreted. If the feature contains multilingual content it is recommended that this value remains undefined or is set to the dominant language of the feature. + +##### The _encoding_ object defines the text encoding within the _value_ objects. Values should be one of "ISO-8859-1" (See [ISO/IEC 8859-1 [11]](https://www.iso.org/standard/28245.html)) or "UTF-8" (See [RFC3629 [12]](https://datatracker.ietf.org/doc/html/rfc3629)). This value is optional and remains undefined if not included. ##### + +##### The _token-schema_ object is included to allow reference to standard sets of token values or symbols for certain _features_ with specific _mime-types_. For example, it could be used to specify which phonetic pronunciation alphabet is present in a pronunciation feature, or refer to a specific standard set of intents for a given domain in a semantic feature. The format of this object is not mandated and is expected to be defined in derived specifications. This value is optional and remains undefined if not included. Figure 2 shows one example of how the _token-schema_ might be used to reference a specific text tokenizer.##### + +##### Figure 4 shows an example of how _token-schema_ might be used to reference a standard semantic schema. Both examples are illustrative, not normative. ##### + + +![Screenshot 2023-06-15 at 6 11 51 PM](https://github.com/NSouthernLF/docs/assets/101130471/083c49b1-4a93-4c98-aef1-fd471d9a95b3) +##### _Figure 3. A feature object containing tokens objects with spans._ ##### + +##### Each token object can also optionally contain a _span_ object. This indicates the time span of this particular token. As described in section [1.3 Span](https://www.iso.org/standard/28245.html), spans can be defined as absolute times or relative offsets. In a token object _start-offset_ and _end-offset_ are relative to the _start-time_ of the parent dialog event (or the implied dialog event start time defined by its _start-offset_).##### + +##### Figure 3 shows an example where the tokenized words of the text feature for Figure 2 are each given an individual _span._ ##### + +### 1.5 Confidence, Linking, and Stand-Off Annotation +##### Each feature should represent just one aspect of the dialog event. Each token object within each feature can be linked to another token object in the event using the object links. ##### + + +![Figure4](https://github.com/NSouthernLF/docs/assets/101130471/20c637f9-28ec-48e8-b17f-c9abfdc3e6be) +##### _Figure 4. A simple example of linking between features with confidence assigned to a feature._ ##### + + +##### Figure 4 shows an example of a semantic feature which links to the text that it represents. ##### + +##### The optional _links_ object is an array of references to the _value_ (or parts of the _value_) of other token objects using a JSON Path with some extensions. There is no restriction on how many links a feature can have. ##### + +##### Unlike the text in Figure 2, in this example, the text is not tokenized. The semantic token with the name _"intent"_ is linked to the whole text of the utterance as defined by the JSON path ```$.my-text-feature.tokens[0].value.``` ##### The semantic token with the name _"date"_ however is linked to a sub-string of the source text - the word _tomorrow_ - via the JSON Path ```$.my-text-feature.tokens[0].value.substring (34,41)```. Section 1.6 describes how to use JSON Path references in more detail. +##### The optional _confidence_ object is a number carrying a measure of the confidence that the information contained in the associated value is ‘correct’. Confidence values are expressed as a probability (real number between 0.0 and 1.0). Derivative specifications may override this range for specialized situations. + +##### Recall that the format of a given _value_ object is defined by the _mime-type_ and will vary between features. ##### + +##### Feature layering and cross-referencing are both important parts of the dialog event standard. Together they permit the different features of an utterance or linguistic event to be kept separate but also linked logically and temporally with each other. This approach is termed stand-off-annotation.[9] ##### + +### 1.6 JSON Path and the substring() extension + +##### Individual elements of the linked object are defined as strings containing JSON Paths. JSON Path is a notation for querying and manipulating data stored in JavaScript Object Notation (JSON). +##### At the time of publication, there is not an agreed standard definition for JSON Path. There are a number of implementations based on [JSON Path [7]](https://goessner.net/articles/JsonPath/) that are broadly compatible. There is also an IETF task force [RFC6901 [8]](https://www.rfc-editor.org/rfc/rfc6901) working on a common standard.##### + +##### JSON Path expressions in _links_ objects are cross-references from one token object to the _value_, or part of a value, of another token object. For these references, the root document (designated by ‘$’) is the _features_ object in the current dialog event. There is currently no way to specify a link between features that are not contained in the same dialog event. ##### + +##### Implementations should support one additional non-standard extension to JSON Path - namely the _substring()_ function. This extension is added to allow token objects to reference a specific substring within another token object, for example, a word, phrase, or arbitrary sequence of characters. ##### + +##### The _substring_ function can be invoked at the tail end of a path. The input to this function is the output of the preceding path expression. The substring function takes two ordered comma-separated parameters, the index of the start and end character in the substring. The first character of the string is at the index _zero_. ##### + +##### For example the JSON Path: ```$.my-text-feature.tokens[0].value.substring(34,41)``` will return the 34th to the 41st character of string contained within the object referred to by the JSON Path ```$.my-text-feature.tokens[0].value```. If the referred object is not a string, or the substring is out-of-bounds then the resulting behavior is undefined. ##### + +##### Token objects can contain content that is specified externally via a _value-url_ string. If this external content is expressed in JSON format then the JSON path should be interpreted as if the _value_ had been specified locally. This is functionally equivalent to reading the content of the document referenced by value-url and placing it in the equivalent local _value_ object before applying the links reference. ##### + +### 1.7 Alternates + + ![Diagram5](https://github.com/NSouthernLF/docs/assets/101130471/b878d7c6-690c-4046-a73e-6a0cd76dd8e9) + +##### _Figure 5. Defining alternate values for a feature._ ##### + +##### Spoken or written language rarely has a single unambiguous interpretation. For this reason, the dialog event object allows for multiple interpretations of a given event. As has already been seen, the _tokens_ object is used to represent the preferred interpretation of a given feature. In addition to this, the optional alternates array can be used to express alternate interpretations of this feature. For example, in Figure 5 it can be seen that the preferred interpretation of the feature named _my-text-feature_ has the _value_ “what is the weather forecast for tomorrow” and a _confidence_ of 0.96. The _alternates_ array contains one additional interpretation of the utterance with the _value_ “what is the weather forecast for thursday” and a _confidence_ of 0.73. ##### + +##### The _alternates_ object is an array of an array. Each object in the outer array represents an alternative interpretation and each inner array represents the token objects for this interpretation. These inner token objects have the same format as array elements in the _tokens_ object. ##### + +##### The outer array can contain zero or more interpretations. An empty _alternates_ array implies that there are no alternate interpretations and is equivalent to the _alternates_ object not being present in the event. ##### + +### 1.8 Nomenclature +##### This specification uses ‘kebab-case’ (i.e. hyphenated lowercase) for all nominal property names, for example, _start-offset_ and _value-url._ User-defined object keys such as the dictionary names used in the _features_ section can use any valid JSON expression. It is recommended but not mandated that any normative key names in derived specifications follow the same kebab-case convention. ##### + + +### Chapter 2. Schema +##### The structure of a JSON dialog event is defined below as a JSON Schema. ##### + +![Diagram5](https://github.com/NSouthernLF/docs/assets/101130471/20a4a092-6f06-4f9d-8c0d-2390ba15824d) + + +### Chapter 3. References +* [1] https://www.ecma-international.org/publications-and-standards/standards/ecma-404/ ECMA-404 The JSON data interchange syntax +* [2] https://www.iso.org/iso-8601-date-and-time-format.html ISO 8601 Date and Time Format. +* [3] https://datatracker.ietf.org/doc/html/rfc3339 Newman, Chris; Klyne, Graham (July 2002). Date and Time on the Internet: Timestamps. IETF. doi:10.17487/RFC3339. RFC 3339. Archived +* [4] https://www.rfc-editor.org/rfc/rfc5646.txt RFC 5646. BCP 47. Tags for Identifying Languages. +* [5] https://www.ietf.org/rfc/rfc4646.txt RFC 4646 Regarding Best Practice for Tags for Identifying Languages +* [6] https://datatracker.ietf.org/doc/html/rfc2231 MIME Parameter Value and Encoded Word Extensions: Character Sets, Languages, and Continuations +* [7] https://goessner.net/articles/JsonPath/ JSONPath - XPath for JSON. Stefan Goessner. +* [8] https://www.rfc-editor.org/rfc/rfc6901 RFC6901. JavaScript Object Notation (JSON) Pointer +* [9] Pose-Rodriguez, Javier & Lopez, Patrice & Romany, Laurence. (2014). A Generic Formalism for Encoding Standoff annotations in TEI. +* [10] https://openvoicenetwork.org/docs/2832-2/ Interoperable Dialog Packet Requirement Specification. +* [11] https://www.iso.org/standard/28245.html ISO/IEC 8859-1:1998 Information technology — 8-bit single-byte coded graphic character sets — Part 1: Latin alphabet No. 1 +* [12] https://datatracker.ietf.org/doc/html/rfc3629 RFC3629. UTF-8, a transformation format of ISO 10646 + +### Chapter 4. Glossary of Terms +![graph4](https://github.com/NSouthernLF/docs/assets/101130471/64d076a4-117f-4348-b941-804675d841c1) + +### Chapter 5. Decision Log +##### This section documents some of the key design decisions that were made by the team during the development of this specification. It is informative, not normative. ##### + +![5b](https://github.com/NSouthernLF/docs/assets/101130471/2f1e2256-d826-4fd4-898c-0ca7dd083187) + +## OVON Interoperable Event Restrictions ## + +![LastDiagram](https://github.com/NSouthernLF/docs/assets/101130471/f102b54b-a431-4bb2-b194-5f054b96390a) +