Saturday, August 8, 2026

Koa Update 2: Argument Marking

Continuing to continue our major grammar update!

2. Argument Marking

This is the area where my summer epiphany represents the biggest change to past understanding and usage, as well as the most significant resolution of past woes. I'm both thrilled and almost afraid to announce that we finally have a satisfying, simple, solid system of participant marking, including indefinite participants and their interaction with existentiality.

In a way this is going all the way back to 2003, because I was onto something with my thinking about referentiality that got muddled later on. The muddling was understandable, I think, given my own language background, and intuitive familiarity with languages in which specifiers mush together specificity, definiteness and existential quantification.

The important foundational question is why Koa requires argument marking at all. It's not because definiteness simply must be granularly specified for every participant in order for language to be interpretable, because...Polish, Latin, Malay, Korean, etc. No, arguments have to be marked in Koa because the language is omnipredicative: that marking is the only way we can recover syntactic function in a world where every lexeme can be a noun OR a verb OR an adjective. The marking doesn't have to be a perfect taxonomy of referentiality and discourse stage relationship, it just has to say "hi, this is an argument."

Koa's new system was catalyzed when I recalled that the Samoan article system uses its "the" to mark participants that are referential rather than definite. Whereas English would say "I spent my day lying in bed reading a book" because the book in question is not yet on the discourse stage, Samoan would say "the book" because the speaker is referring to a real, specific book out in the world. It's an elegant system that looks similar to English articles at first glance, but is really quite different.

The most critical fork for Koa now hinges on that same question, which could be put like this: Does the truth value of the proposition depend on the identity of the instance? In other words, is the identity of the referent anchored to a specific individual in the mind of the speaker? Or do there exist many possible referents that could could satisfy the proposition?

This is a spectrum, not a binary. At one pole we have extremely specific, defined referents, as in "my wife is vacationing in Yakutsk"; at the other pole are completely unanchored referents as in "I would like a glass of eau de vie, please." In the middle is a gray area in which individual speakers' judgments may vary. For example, taking a statement like "a mole made a mound in the middle of my prize begonia bed," the speaker could follow this up with

"...I guess I should lay chicken wire under my flower beds next year," OR
"...and when I find it, it will answer for its crimes."

Of course, there was in fact one specific mole that caused the mound under discussion. The relevant information, though, is how the speaker is framing the situation in their mind. In one case the speaker's attention is on the existence of moles in general, unanchored to individuals; in the other, the speaker is focused on the specific -- though as yet unknown -- small mammal against whom they have just conceived a vendetta.

My own personal growth at this point is in the recognition that that gray area is not in fact a design problem I need to spend another 27 years seething about: it's a normal, inevitable scenario in human language use. It's okay that speakers have to make a judgment, that that judgment sometimes depends on inscrutable intuition, and that another speaker might make a different judgment in the same situation.

This sensitivity to the anchored↔︎existential continuum is foundational to Koa's updated argument marking system. Arguments must appear, as mentioned above, with some kind of typing particle or structure:

- a specifier or quantifier
- a relator (which outputs as an adjunctive expression)
- a quantificational expression

Note: Where the theme (technically second variable) of a predication is generic or its specificity is irrelevant or discourse-backgrounded, and the predication as a whole is judged to be a culturally conventional state or event, that predicate may appear without marking: luke tusi "read a book, read books." This is theme incorporation, a valence-reducing operation that satisfies the second argument position in the nuclear expression, and the predicate in question is technically a modifier (see §4.3).

2.1 Specifiers & Quantifiers

This is it, folks -- here are the details of the new system.

Specifiers establish how an argument is anchored, identified, presented, retrieved, named, or questioned within the relevant discourse model. They include:

- ka marks anchored specific referents, agnostic about relationship to the discourse stage. The speaker presents the argument as a particular contextually anchored instance, regardless of whether the addressee can identify it. Note that this no longer has anything directly to do with definiteness! Though "the" may often be the most natural English translation, ka does not operate on the same semantic axis as the English definite article.

- ti marks either deictic proximity or presentational/cataphoric reference. It directs attention toward an entity being presented now; it may therefore refer to something physically or conceptually close to the speaker or current discourse point ("this book here in front of me") or introduce a referent that will be important in later discourse ("this guy walks into a bar").

- to marks either deictic distance, or anaphoric retrieval (preexistence on the discourse stage). It directs attention toward an entity already situated away from the current point of introduction; as such, in addition to representing "that" in a distal sense, it is also the closest Koa now comes to an actual definite article: "the referent already present on the discourse stage."

le marks names (§1.2.6).

ke marks an open interrogative argument whose identity is being requested.

Quantifiers, on the other hand, determine the extent to which a predicate’s domain is instantiated or selected. Unlike specifiers, they do not anchor the argument to a particular referent. Quantifiers and specifiers may occur together (see §2.2 below), as they're operating on different semantic dimensions. These include:

- a, the existential quantifier, introduces one or more unanchored instances satisfying the predicate. The speaker suggests that there is, may be, or must be at least one instance satisfying this predicate without selecting a particular one. Again, this is no longer about indefiniteness, though it may sometimes coincide with the use of the English indefinite article.

- po, the universal quantifier, selects the whole relevant extension of a predicate or presents it as a class, kind or generic domain. It refers to the instantiation of the predicate generically, or to all members of a contextually restricted set.

- co, the unrestricted or free-choice quantifier, selects any arbitrarily chosen member of the relevant domain: the identity of the selected member explicitly does not matter.

The places where this may feel most counterintuitive to a person with an existing Koa background (i.e. probably only me and Robert) will be in situations like

with a mass noun:
a vua i pahu la lavi ka iku "(some) rain fell through the window"

a with plurals:
ni mupahu a kina la kuvu "I dropped (some) pens on the ground"

ka with indefinite-feeling referents:
ka mina ni i soi me simi "A friend of mine called on the phone"

It's going to take some getting used to for sure. But -- now, kind of amazingly, and for the very first time, it's easy and clear to understand exactly what each particle really means! There are judgment calls, but not definitional confusion, and no unnecessarily granular distinctions. AND we've escaped the chronicle of indefinite referent miseries...by simply stepping away. It turns out that indefiniteness qua indefiniteness is not really that important to mark explicitly after all.

2.2 Operator Nesting / Scope layering

There is a bit more to the show here, because specifiers and quantifiers can also stack. When a quantifier scopes an argumental expression already marked by a specifier, it takes on a more literal quantificational meaning:

a ka sona "some of the ducks"
a ti sona "some of these ducks"
a to sona "some of those ducks"

po ka sona "all of the ducks"
po ti sona "all of these ducks"
po to sona "all of those ducks"

If we don't want to anchor "some" and "all" to a contextual set -- just "some ducks" (out of all the ducks there are) or "all ducks" (period) -- there are a couple options. We can use a and po on their own, for light quantification:

a sona i miku "some ducks are complicated / there are complicated ducks"
po sona i miku "ducks are complicated (as a class)"

...but if we really want to lean into the quantification, we'll need help from an additional lexical predicate: visi "all, every" and ioko "a set, some" (and see §2.6 for further discussion). These can either follow the argument as a modifier, or become a quantificational expression with pi:

a sona ioko i miku "there are some ducks that are complicated"
ioko pi sona i miku "some ducks are complicated"

po sona visi i miku "every duck is complicated"
visi pi sona i miku "all ducks are complicated"

The semantic difference between the options in each case is very small. If we really drill down, the expressions with a and po are focusing attention more on the predication (i.e. being complicated), whereas the expressions with ioko pi and visi pi are focused on the quantity: some ducks, all ducks. In actual usage, though, these essentially represent equivalent alternatives.

Of course, there are also basic quantificational expressions like the following:

poli pi sona
 "a lot of ducks"
aiva pi sona "quite a few ducks"
nai pi sona "as unspecified number of ducks"
neso pi sona "some smallish number of ducks"
tele pi sona "several ducks"
vaha pi sona "not many ducks, few ducks"

toa pi sona "that many ducks"
kea pi sona? "how many ducks?"

See §6.1 for discussion of interactions between relators and argument marking.

2.3 Compound Proforms (Correlatives)

Argument markers also join together to create compound pronominal forms analogous to the correlatives of Esperanto. The first component of these forms indicates deixis, quantification or interrogative status, and the second distinguishes unanchored from anchored/human reference. These are somewhat lexicalized and -- clearly -- not 100% reflective of the individual meanings of the components.

unanchored anchored/human
what/which kea “what” keka “who/which one”
this tia “this, all this” tika “this one”
that toa “that, all that” toka “that one”
some aha, aa “something” aka “someone/some of them”
any coa “anything” coka “anyone/any of them”
all poa “everything” poka “everyone/all of them”
none naha, naa “nothing” naka “no one/none of them”

Sadly, our long-standing correlatives with hu (hua "something," nahua "nothing," etc.) are no more. I like hua rather more than aha, but as it's probably ridiculous to retain it as an alternative, I've found a new use for it (see §2.6!).

The other class of compound proforms are pronominal predicates, formed by reduplicating the corresponding pronominal particle:

nini "I"
sese "you"
tata "he/she/it"
nunu "we"
soso "you all"
tutu "they"

Pronominal predicates deserve a longer discussion than I'm going to give them here. Essentially, they are used either when a nuclear function is needed syntactically (tika i nini "this one is me," e.g. in a photograph) or for stress, emphasis or contrast (ana ta la nini, na la tata "give it to me, not to him").

As mentioned previously in §1.2.5, all these forms are exceptional in being licensed for argumental distribution without an argument marker. The only circumstances in which they might appear with a specifier or quantifier might be e.g.

oo he ko ni sano to aha -- halu ko lolo ka molo se
"oh when I say that something -- I wanna hold your hand"

ai se aima ka pacoi ni? ia, ka sese i ce koa i taa ka nini
"do you like my drawing? yes, yours is much better than mine"

2.4 Existential Structures

Along with "indefinite participants," expressing existence has historically been another of the abject intractables. Part of this was due to gaps in the description/understanding of the argument marking system that made forms with a feel insufficient. In general, though, I've also had very low tolerance for there being more than one way to do anything unless the selection criteria were minutely defined.  This blind spot now feels quite quaint -- multiple overlapping syntactic options obviously need not bespeak a design failure.

Existence in Koa, then, can be expressed in any of three ways: through existential quantification with a, with the predicate tai "exist, obtain," or via a comitative structure with me. The differences between these three strategies are highly discourse-sensitive.

With a (and compound proforms with a- as the first component), the least marked strategy, the existential force is present but neutral, with discourse orientation towards the figure under discussion:

ai a vitu i ne ka noni?
"is a dragon / are (some) dragons in the mountains?"
"is there a dragon / are there dragons in the mountains?"

ni halu ko sano se aha
"I want to tell you something"
"there's something I want to tell you"

Other quantifiers produce related existential interpretations with a locative nucleus. po takes the predicate generically or as a class, while co, in interrogative and negative contexts, gives an unrestricted "any at all" sense.

ai po vitu i ne ka noni?
"are (there) dragons (as a class) in the mountains?"
"do dragons occur in the mountains?"

ai co vitu i ne ka noni?
"are (there) any dragons (at all) in the mountains?"

Using tai as the clause nucleus makes existence explicit and at issue. Again all three quantifiers may appear, with different semantics:

ai a vitu i tai ne ka noni?
"does a dragon / do (some) dragons exist in the mountains?"
"is there a dragon / are there (some) dragons in the mountains?"

ai po vitu i tai ne ka noni?
"do dragons (as a class) exist in the mountains?"
"do there exist dragons in the mountains?"
"do dragons occur in the mountains?"

ai co vitu i tai ne ka noni?
"do any dragons at all exist/occur in the mountains?"
"are there any dragons at all in the mountains?"

Lastly, (i) me, literally "be with, have" is more setting-oriented or presentational:

ai (i) me vitu ne ka noni?
"are there dragons in the mountains (that we're talking about)?"
"in the mountains, are there dragons?"

po may optionally quantify the existential argument, emphasizing a generic reading; co is also possible:

ai (i) me po vitu ne ka noni?
"are there dragons (as a class) in the mountains (under discussion)?"

ai (i) me co vitu ne ka noni?
"are there any dragons (at all) in those mountains?"

These structures with me can also be rephrased with the setting as an overt first argument, in contexts where that setting is particularly topical:

ai ka noni i me vitu?
"do the mountains (that we're talking about) have dragons?"
"are there dragons in those mountains (that we're talking about)?"

2.5 Negative Existentials & Concord

Here's another imponderable that we can resolve at last. In general, the unmarked preference is for negative polarity to be established as early as possible in the clause; after this, negative concord is accepted but not required:

Without concord:
naka i tule "no one came/there was no one who came"
na a vitu i ne ka noni "there are no dragons in the mountains"
ni na ana aka aha "I didn't give anybody anything"

With concord:
naka i na tule "no one came/there was no one who came"
na a vitu i na ne ka noni "there are no dragons in the mountains"
ni na ana naka naha "I didn't give anybody anything"

I hesitate to say that expressions with negative concord may be read as somewhat more emphatic than those without. That would be reasonable, but in a language intended to be used by speakers from any language background, I don't want to force that interpretation.

Negative force arriving via a later negative constituent is possible, and interpretable, but would be very marked:

Neutral: ni na sano a mea mono "I didn't say a single thing"
Marked: ni sano na a mea mono "I said not a single thing"

One thing to call out for historical Koanists (Koaists? Koalogists??) is the negative quantifier na a, previously nahu. I waver on whether the two morphemes can be written together since naa also exists independently with the meaning "nothing." Written spacing conventions have also clearly gone back to the drawing board, though -- see §7 -- so we'll take this on definitively a little later.

2.6 Disambiguation & Emphasis

In addition to the argument markers discussed so far that provide the structurally scaffolding semantics of quantification, referentiality and discourse operations, additional emphasis and disambiguation in these areas are available via lexical predicates.

Quantification

Here the qualifying predicate may either follow the argument in question as a modifier, precede it in a quantificational expression or, when a postnuclear argument is quantified, may follow the clausal arguments as a serial verb (see §3.3):

ioko "some; a set"
a asi ioko i vihe OR ioko pi asi i vihe
 "some ideas are green"

visi "all, every"
po kunu visi i va mene la ka lani OR visi pi kunu i va mene la ka lani
"all dogs go to heaven"

sai "whole, entire"
ai se ia suo ka kulo sai? "did you really eat the whole bowl?" OR
ai se ia suo sai pi ka kulo? "did you really eat all of the bowl?" OR
ai si ia suo ka kulo i sai? "did you really eat the bowl in its entirety?"

hetu "any, at all"
ai se suso co iki hetu? OR
ai se suso hetu pi iki? OR
ai se suso co iki i hetu?
"did you kiss any frogs at all?"

Discourse

hua "some, a certain" can either reinforce an introduction to the discourse stage when modifying a relational argument with ti or ka:

vo ka/ti kúnika hua i si asu ne...
"once there lived a certain king..." lit. "behold a/this certain king lived there"

manu "present, aforementioned, that we are talking about" reinforces the retrieval of an existing referent from the discourse stage:

ala ea lai la ka kepe manu
"but let us return to the topic at hand"

Specificity

aha "of some kind" lit. "something" and again hua "some, a certain" reinforce nonspecificity with the existential quantifier a:

a hili aha/hua i li si iune ka sie nu visi
"some mouse seems to have stolen all our cheese"

litu "particular, specific, the very" reinforces specificity:

vo ka kane litu i tule, u nu ma puhu
"here comes the very man we were speaking of"

laca "general" reinforces a generic interpretation:

po vitu laca i asu ne to noni
"dragons do generally dwell in those mountains"

Number

Since Koa arguments are not morphologically marked for plurality, overt number can be emphasized by mono "only, singular" or ena "one" for singular, or eli "various, several" for plural:

ai se nae a kúhuto ne ka pese?
"do you see an owlet/some owlets in the nest?"

a kúhuto ena/mono i ne ka pese
"there is one/a single owlet in the nest"

a kúhuto eli i ne ka pese
"there are several/plural owlets in the nest"

...and that wraps up Section 2! Next up will be nuclei and clauses, which will let us close the books on complement clauses at long last. Stay tuned...

Monday, August 3, 2026

Koa Update 1.2: The Predicate

Continuing our major summer update from where we left off...

1.2 The Predicate

Koa predicates contribute lexical meaning. Other than proper names (see §1.2.6), primitive predicates (roots) are all bisyllabic, of the orthographic form (C)VCV, CVV, VVCV and very occasionally VVV. Detailed phonological constraints for predicates are described here.

A minimal description of each predicate root consists of its lexical meaning, lexical aspect (§1.2.1) and canonical argument frame (§1.2.2).

1.2.1 Lexical Aspect

Every Koa predicate is lexically stative or eventive. I had recognized this to be so for almost 15 years, but this is my first attempt to articulate the criteria of the system. Note that this is a semantic lexical distinction, rather than e.g. a noun-verb distinction, and a larger derived expression may have a different compositional aspect class than its root(s) (see §1.2.1.2 below).

1.2.1.1 Statives & Eventives

A stative predicate describes, broadly, the way something is. It presents a condition, property, identity, relation, perception, attitude, or classification as simply holding of its participants. Examples:

Classification: sene "be a cat," talo "be a house"
Identity: ka mama ni "be my mother," le Iúliki "be Julie"
Property: puna "be red," cahi "be sincere"
State/condition: loe "be cold," nuku "be asleep"
Attitude: tulu "be angry," ihi "miss, long for"
Relation: imi "be the same entity," ilo "know," oma "belong to"
Experience/perception: nae "see," loha "love," halu "want"

An eventive predicate presents something as happening: an action, process, transition, or other occurrence unfolding or taking place in time. Examples:

Activity: colo "run," puhu "speak," hana "work"
Process: sunu "grow," kuno "melt"
Transition: ata "arrive," lilo "become," mimua "die"
Bounded event: tui "break, smash," hau "build"
Momentary occurrence: lou "hit," hupa "jump"

Note that the stative/eventive distinction is not based in agency, transitivity, durativity, telicity or perfectiveness; these are separate dimensions encoded in different ways.

1.2.1.2 Aspectual Coercion

Some particles and structures have the effect of coercing their scope into a particular lexical aspect, based on the inherent meaning of the output. This change can happen in either direction.

Stative -> Eventive
kulu "hear" -> makulu "be listening to" (experience -> activity)
mua "be dead" -> mimua "die" (state -> transition)
nae "see" -> nae élakuva "see a movie" (experience -> bounded event)

Eventive -> Stative
colo "run" -> tecolo "can run" (activity -> property)
suo "eat" -> vasuo "habitually/generally eat" (activity -> property)

1.2.1.3 Interaction with Grammatical Tense & Aspect

As I first remarked in 2012, the verbal equivalents of statives and eventives in Creoles tend to interact differently with tense and aspect. On the grounds that this represents a graceful economy likely connected with human language universals, I adopted a similar system for Koa as well. Essentially, the default interpretation of the bare nucleus, and the nucleus with anterior si, is different for statives and eventives:

Bare Nucleus Anterior
Stative Present Past
Eventive Perfective Past Pluperfect

Another way to say this is that eventives are interpreted by default to overlap the present reference time, whereas statives are interpreted to have been completed before it.

Thus:

ni loha "I love" (present)
ni si loha "I loved" (past)
ni mene "I went" (perfective past)
ni si mene "I had gone" (past before point of reference)

However, note! An eventive with progressive marking (ma) patterns with statives in temporal/aspectual reading. In this respect, the default tense/aspect interpretation of predicates is sensitive to perfective/imperfective status.

ni ma mene "I am going" (present continuous)
ni si ma mene "I was going" (past continuous)

That said, past continuous "I was going" is an equally reasonable interpretation in context of ni ma mene, as is pluperfect continuous "I had been going" of ni si ma mene. I should emphasize strongly that although the above tense and aspect interpretations are the least marked, others are often entirely possible; indeed, ni mene by itself could have any of these interpretations in contexts where TAM specificity is either not needed or has been provided elsewhere or by other means.

1.2.2 Semantic Frames

Perhaps it is already obvious given the fact that lexical roots in Koa are called "predicates," but the foundation of this language owes a lot to symbolic logic. Despite this underpinning, I actually think the way the system plays out in practice is very intuitive, clear and easy, in a way that e.g. Loglan unfortunately is not. In usage, it doesn't feel like a "logical language," just a language in which things that are usually left to guesswork/intuition/experience are clearly articulated in advance.

Part of this is in the specification of what each predicate "means." For example, it is not enough only to say that nae means "see." We need to know that it is stative (see above), that it licenses two canonical arguments, and that these -- in order -- are assigned the frame roles (see below) of experiencer and theme. Put more technically:

Each Koa predicate has a lexically specified semantic frame comprising its lexical meaning, canonical arguments, and the frame roles of those arguments.

This is one of the most important, basic things to understand about Koa, yet something I don't think I've ever actually stated explicitly.

The "frame role" is loosely the semantic role of the argument with respect to the predicate, but we need not specify the roles with granular precision: since much of this information is implied by the semantics of the predicate itself, the roles need only distinguish feasible participants such that it is clear what is going on. The frame roles currently defined for Koa are:

Initiator (I) - agent, force
Experiencer (E) - experiencer, perceiver, emoter, cognizer
Theme (T) - theme, patient, undergoer, bearer, percept, content, situation
Goal (G) - goal, recipient, beneficiary, destination, target
Reference (R) - relatum, standard, comparator, reference point, landmark
Causer (C) - causative agent, external causer

A dictionary entry for nae, then, might look something like this:

nae S2 see ⟨E,T⟩

1.2.2.1 Frame Types

As of now, the following semantic frame types are known; in the following notation, "S" stands for "stative," "E" for "eventive," followed by the number of canonical arguments licensed by the predicate.

S1: tai "exist," ava "be open," nahi "be a crab"
S2: loha "love," oma "belong to"

E1: hupa "jump," kuno "melt"
E2: lilo "become," tui "smash"
E3: ana "give," kumu "teach"

Terms for the number of a predicate's canonically licensed arguments are: "monovalent" or "one-place"; "bivalent" or "two place"; and "trivalent" or "three-place."

Initiators and experiencers tend towards argument 1, with themes, goals and references expressed with later arguments; with trivalent (E3) predicates, the licensed post-nuclear order is goal-theme (see §5.3). One could say -- not with perfect accuracy, but perhaps with useful explanatory force -- that Argument 1 tends to represent the subject and Argument 2 the object -- or with ditransitive (i.e. trivalent) predicates, Argument 2 the indirect object and Argument 3 the direct object.

1.2.2.2 Alternative Frames

In some cases, because fluidity and ease of use is more important than being precious about the logic of prescribed semantic frames, and because so much is easily recoverable from context, a predicate may be used with an alternative frame. Some examples:

Canonical: ava S1 be open ⟨T⟩
Alternative: ava E2 open ⟨I,T⟩

Canonically, "I opened the door" would translate to Koa ni muava ka ovi, mu being the causative particle: literally "I caused the door to be open." There is nothing to be lost, however, in allowing ni ava ka ovi with the same meaning. Similarly,

Canonical: lopu S1 end, come to an end ⟨T⟩
Alternative: lopu E2 end, bring to an end ⟨I,T⟩

Thus, for "we ended the meeting," we could have either the canonical nu mulopu (causative) ka níkete or alternative ni lopu ka níkete.

Lastly, perhaps we might see:

Canonical: olo E2 smell ⟨E,T⟩
Alternative: olo S1 smell ⟨T⟩

"The flower smells sweet" would usually be ka lulu i paolo mo make, literally "the rose is smelled like a sweet one," but conceivably could appear as ka lulu i olo mo make, lit. "the rose smells like a sweet one."

Ideally, this being (at least theoretically) an IAL, the lexicon should list the canonical frame along with any other known and accepted alternative frames to make usage options and interpretation clear and explicit.

1.2.2.3 Argument Omission

Arguments may frequently be omitted under certain circumstances: if their referents are recoverable from the discourse context, irrelevant, or generic. For example, in response to the question "Did you finish reading that book," we could encounter any of the following responses:

ia, ni io luke ta "yes, I already read it"
ia, ni io luke "yes, I already read Ø"
ia, io luke ta "yes, Ø already read it"
ia, io luke "yes, Ø already read Ø

In some contexts there is no clear overt referent for a given argument in the first place, in which case omission is standard:

i loe netia "[it's] cold in here"
i mevua hetana "[it's] raining today"

Note that these omissions do not constitute alternative argument frames, as the argument slots and their semantic roles are still present underlyingly.

1.2.3 Expression Functions

As discussed somewhat in §1.1.1, rather than each predicate having a lexically specified distribution class like in IE languages, syntactic role in Koa is assigned constructionally. This is to say: the context of each predicate or predicate expression in syntax determines its function and how its meaning should be construed.

I had been using the term "predicate function" to describe this concept, which I still think is rather snappier. That terminology obscures, however, the crucial fact that the same constructional criteria apply equally to individual monomorphemic predicate roots like kuhu "owl," to polymorphemic compounds like talokuhu "owl house" or kúhuto "owlet," to specified expressions like ka kolu ni "my school," or even to whole clauses like to kúhuto i asu ne talokuhu ne puu ne uko ka kolu ni "the owlet under discussion lives in an owl house in the tree next to my school." "Syntactic function" is another possibility, but this feels too broad in the sense that it doesn't anchor the domain of description to the predicates which are at the core of it.

Six expression functions are currently described for Koa. The terminology for these functions constitutes a major change from tradition, being an almost entirely new set of terms intended to communicate clearly what each function is doing rather than which IE lexical class it is somewhat reminiscent of. These functions include:

Argument: These supply the participants in a predication and are licensed by argument markers or the argumentizer ko. Previously referred to as "nouns" or "nominals."

Nucleus: The semantic and syntactic center of a clause; it expresses the state, event, property, classification, identity, location, or relation being predicated and organizes the interpretation of the clause’s arguments. I had been using "predicator" for this purpose until earlier today, but this was desperately easy to confuse with "predicate" and I was chagrined to discover I had inadvertently been doing so with regularity. Nuclei are licensed either by nucleus markers or by pronominal particles (or occasionally Ø in casual speech contexts with clear recoverability from discourse). Previously "verbs" or "verbals."

Modifier: Restricts, characterizes, classifies, or elaborates another predicate or predicate expression; normally follows the expression it modifies. Bare predicates provide unmarked modification; argument-marked expressions provide referential, often possessive, modification; and clauses are licensed for modificational function by the clausalizer u. Previously "adjectives" or "adjectivals."

Adjunct: Supplies a circumstance under which a larger predication is to be interpreted, ordinarily without filling one of that predication’s lexically specified core participant roles. Licensed by relators or by the adjunctive clausalizers ve or ha. Previously referred to as "adverbials" or other loosely defined terms.

Quantificational Expression: Delimits the number, amount, measure or extent of a domain; licensed by the quantificationalizer pi. This function was understood tenuously in the past, but needs more rigorous description.

Discursive: Manages, relates, evaluates, repairs, or advances the discourse itself rather than serving as part of the grammatical scaffolding of a clause. These typically appear as an unmarked predicate in between clauses. Despite their having come into wide and discourse-critical use, I had not recognized them as a new function until last week!

1.2.4 Proper Names

Proper names in Koa constitute a special type of predicate. Unlike lexical predicates these are an open class, the only restriction being -- ideally -- that they conform to Koa's phoneme inventory and (C)V syllable constraints. (Letters and spellings from other languages may be permissible, especially in writing, but may not be pronounceable!) Stress should be marked if not penultimate.

Names are licensed by the argument marker le when integrated into syntax; le is not used in direct address.

There was some talk once upon a time of le potentially having wider use in mentioning or quoting...this will need to wait a little longer for full consideration.

1.2.5 Pronominals

Pronominal predicates are formed compositionally and are unique among predicates in being licensed for argumental function without argument markers. Thus naka i ilo toa "no one knows that," not *po naka i ilo ka toa.

These fall into the two general classes of correlatives (toa "that," kea "what," aka "someone," naa "nothing," etc.) and personal pronouns (nini "I," sese "you," etc.). See §2.3 (upcoming) for full discussion.

1.2.6 Suffixation

A discussion of predicates in Koa would not be complete without mentioning suffixation. This has existed as a morphological process almost since the beginning, but without much formal description; unfortunately, I'm still not quite prepared to give this area the treatment it deserves. This section represents the show so far.

Koa suffixes mainly encode derivational operations: Iúli-ki "Julie" (diminutive), cáno-a "sir/madam" (honorific), vúa-mi "raindrop" (distinct part of a whole), sáhi-lo "bar" (place of/where). Of those which appear to encode inflection, it is sometimes unclear whether they are genuinely suffixes or simply postposed particles; this is particularly so with directionals which may bear stress. Thus ka séne-ni "my cat"; mene(-)ó "go away."

As suffixes extend the length of a phonological word rightwards, a written accent is needed to maintain the stress on the original syllable of the predicate.

A predicate may bear more than one suffix, layering scope leftwards: ka sáhi-lo-ki-ni "my cute little bar." However, a suffix scopes only the predicate to which it is affixed, and cannot apply to larger structures of constituents:*

ka [séne]-ni sopo "my cute cat," but not *ka [séne sopo]-ni
[súo]-te u ni ipo lia pi sahi "meal (instance of eating) where I drank too much wine," but not *[súo u ni ipo lia pi sahi]-te

Note that suffixes do not exhibit functional coercion, only adjustment to the semantics of the predicate concerned. The external syntactic function of a suffixed predicate, or expression containing a suffixed predicate, is determined by the usual constructional criteria. However, it may certainly affect lexical aspect and valence: suo "eat" is a two-place eventive (E2), but súote "meal" is a one-place stative (S1).

*This is one of the restrictions on Koa's core tenet of universal applicability, which I didn't write about in the previous post but perhaps should have. Essentially: A grammatical operation available to a given domain is presumed available to all expressions in that domain unless a restriction is explicitly established. Thus anything you can do with/to one argumental expression (like ka puu "the tree") can also be done to any other argumental expression (like ka puu u si tuu ne pole ka iku-ni he kei ka tóto-te-ni "the tree that stood outside my window during my childhood"). A bookmark for the reference grammar someday.

Saturday, August 1, 2026

Koa Update 1: Lexical Class & Particles

Up early on a Saturday, the rest of the house still sleeping, time for Koa? Definitely.

I've been panicking a bit, frankly, at the scope of my table of contents, which has continued to grow over the past week. I just realized with relief, though, that for now at least I don't need to write a full description of everything! Initially I think I just need to identify the things that have changed from the understandings and assumptions that have been demonstrating on this blog. And so...

1. Lexical Class in Koa

Koa entirely lacks lexical class (nouns, verbs, adjectives, adverbs) in the way familiar to European languages, predicate function being determined constructionally from syntactic context. However, Koa does have a different type of lexical class in the form of the distinction between particles and predicates. This is not new, but it is worth saying overtly.

Koa being a designed, planned language, all lexical classes are closed except for those predicates which are proper names. However, new meanings can be created via derivation, compounding, etc.

1.1 The Particle

Koa particles contribute grammatical (syntactic, morphological, derivational, pragmatic) meaning. They are monosyllabic in form, most typically (C)V: a, ki, te, etc. Some are orthographically VV, but still represent phonological monosyllables: ai, io.

Particles may frequently be compounded, in which case the scope of each leftward particle layers over the material to its right. When scoping a predicate, particles are typically spoken as a single accentuation group with that predicate: that is, as a single phonological word. When occurring independently they are accented on the last member unless that compound has been accepted as an independent predicate root: haná "unless" lit. "if not," civé "although" lit. "even [given] that"; but meno "regardless, nevertheless, either way" lit. "with without," nini 1sg emphatic lit. "I I."

1.1.1 Typing & Distribution

Particles have both a typing function and a distributional outcome. Their typing function determines how the material within their scope is construed internally; the completed construction’s distribution determines the syntactic positions in which it may occur. Thus:

ka types its scope as an argument (ka sene "a/the cat") and forms an argumental construction, licensing it for argumental distribution (ka sene i peme "the cat is soft").

i types its scope as a nucleus (i peme "is soft") and forms a clause (ka sene i peme "the cat is soft").

ko types its scope as a clause (ko ka sene i peme "that the cat is soft") and forms an argumental construction (ni lule ko ka sene i peme "I think that the cat is soft").

ne types its scope as an argument (ne miti "on [a/the] bed") and forms an adjunctive construction (ni mahiva ka sene ne miti "I am petting the cat on the bed").

pi types its scope as a quantifier (kume hitu pi "17 [of something]") and forms a quantifier construction (ni mahiva kume hitu pi sene "I am petting 17 cats").

...and so on. This is also not new, but it has never been articulated and made plain before! I've adjusted the particle listing in my lexicon to include this (internal type, external distribution, etc.) and a number of other important attributes.

1.1.2 Particle Categories

This is rather nifty -- I've finally got my particles organized into 15 groupings by type and function, their previously having formed a kind of vague cloud I was never sure how to describe systematically. Particle categories include:

Argument markers, marking constituents being used as -- surprise! -- arguments. These include a subclass of specifiers (ka anchored reference, ti proximal/cataphoric, to distal/anaphoric, ke interrogative, le proper name) and quantifiers (a existential/nonempty, co unrestricted, po universal/generic). You may notice some big business in progress around ka and a, and the absence of hu in this list...this is possibly the biggest deal in my understanding of Koa since 2002 and will be discussed in detail in §1.2.1.1. Note also the absence of initial personal pronouns in this list, allowed for inalienable possession for at least 20 years, but no longer acceptable as argument markers.

Aspectual particles, determining the aspect of predication. We have ca persistive, hu prospective (!!!), io iamitive, ma progressive, mi inchoative, si anterior, su cessative, va habitual. hu in this meaning is experimental, suggesting something that is imminent, about to occur: Le Iúliki i huete cai "Julie is about to make tea." (Don't mind if I do...)

Clausalizers mark the constituent in their scope as a clause, however long it happens to be, and license it to be used in a particular syntactic distribution. ko creates an argument clause, ve a circumstantial adjunct clause, ha a protatic adjunct clause, and u a modifier clause. Yes, I have (hopefully permanently) retracted my elegant Washington Manor idea, to be discussed in full in §6.1.

Emotive particles communicate the emotional state of the speaker. aa understanding/surprise, ee uncertainty, eu disgust, ii pain/dislike/nervousness, oo understanding/confirmation, ui regret/commiseration, uu excitement/pleasure.

Epistemic particles give information about the information: the speaker's commitment to it, its source, or its familiarity status: ho mirative, ku common ground, pu reportative, li presumptive, vu dubitative. Ia, though formally classified as a polarity particle (see below), could also belong to this group to the extent that it represents a full commitment to the information.

Interactional particles interact with the attention of the listener: ei attention-getting, oi feedback-seeking, vo presentative. vo also has a large role in structuring discourse and information.

Junctive particles connect constituents, whether clauses, predicates or even particles (e.g. ia ai na? "yes or no?"). ai interrogative/alternative, au disjunction, e conjunction.

Medial Dependency is a catchall category including particles that form a link between constituents on either side of them, conveying particular information about each. sa, focalizer, places the discourse focus on the constituent to its left, with the gapped clausal residue to its right; pi, quantificationalizer (ugh I know, but I'm really having trouble finding a better term) creates a unit of quantity, formally a quantifier, to its left, which quantifies the domain to its right. In both cases, the material to the right may be elided when recoverable from context.

Modal
 particles contribute mode/mood to the predication: cu irrealis, vi imperative-optative, ki necessitative, lu volitive, te potential.

Polarity particles: na negative, ia affirmative/verum.

Nucleus markers, identifying and typing the nucleus (or nuclear complex) of a clause. These include the general nucleus marker i, and the more specialized ao precative, ea hortative and oe admonitive. It is an open question, probably to be answered by hypothetical future Koa syntactical theorists rather than by me, whether pronominal particles (see below) also fall into this category or simply suppress the appearance of i (i.e. le Iúliki i kiuni "Julie is tired" vs [i?] ta kiuni "she is tired").

Pronominal particles can function as person/number marking of a pronominal first argument of the nucleus, or as independent arguments in their own right: ni 1sg, se 2sg, ta 3sg; nu 1pl, so 2pl, tu 3pl.

Relators create adjuncts, indicating a physical, temporal, conceptual or other relationship with the scoped argument: he temporal, la goal, lo reason, me comitative, mo similative, ne location, no privative, o source, pe reference. Note that (A) ci has been removed as an instrumental marker, this role being taken over by me or i itu for instruments and o for demoted agents; (B) almost all meta-language deriving from IE case systems has been replaced with more neutral terms, e.g. "goal" rather than "dative" or "allative."

Scalar particles indicate the degree or inclusivity of their scope. These include ce degree-raising, iu degree-matching, ie restrictive ("just"), ci inclusive ("even"). In some cases, as with e.g. a ce kane "a real man" lit. "a very man," the predicate's meaning is coerced into a scalar; whether these particles should therefore also be considered to type their scope as modifiers is an open question. ci as an inclusive scalar feels like a much better use than the erstwhile instrumental, and it provides an elegant means for forming concessive clauses (separate post to come).

Lastly, valence particles perform valence operations on the nucleus they scope: hi reflexive/impersonal, mu causative, pa passive.

I'm also pleased to have finally standardized abbreviations for all these, so my glosses will be a little more coherent in the future.

...I was going to try to do the entirety of Section 1 in this post, but there is so very much to say about predicates that I'm going to stop for breath here and pick up at §1.2 next time.

Sunday, July 26, 2026

Excursus for posterity: Syntactic nomenclature

This preamble to a post on purpose clauses was originally written on July 27, 2025. Now that the world has changed this has become rather quaint, but I preserve it here as an artifact of what I was already beginning to feel strongly last year. I note that I was also seeing this in the water as long ago as January 2012, which feels a little poignant to read. I have been struggling with this for such a long time!

There is also an onomastically frustrating pair doing a lot of heavy lifting: on the one hand the nominal clause, referring to a finite clause being used in nominal syntactic positions; and on the other the nominalized clause, a non-finite clause which has formally become a nominal with all the expected syntactic properties. As discussed previously, we really seriously need some better nomenclature here.

As also discussed previously, it's not entirely clear how applicable or useful the language of "finiteness" is for Koa in the first place. I continue to use the terminology out of habit because that's how these structures might be described in other languages, and I haven't come up with anything more Koa-specific. I suppose I could use even more standard language and say "complement clause" instead of "nominal clause," "relative clause" instead of "adjectival clause," and so on, but here's the thing:

In most languages there's a haphazard variety of strategies in use to achieve these communicative needs, and it makes sense to group them under a labeled conceptual umbrella like "complement clause." In Koa, though, most of these are just clauses -- plain old finite clauses, like any other -- and the only difference is which clause marker they each bear. This is analogous to the way that Koa dispenses with formal lexical class, with nominal/adjectival/verbal/adverbial/etc. status being determined entirely by syntactic context. It feels weird to give something a secondary label when its meaning is already transparently derivable from first principles.

Just like all "content words" in Koa are called "predicates," and when I need to refer to their use in a specific context I can say "nominal predicate," or more wordily "a predicate used in a nominal context," I wonder if I could do something similar with clauses. Koa has its own terminology for these things: for example méama, literally "thinger," for the above type of syntactic context.

The mistakes we make

For the last 24 years there have been two areas of Koa syntax that have vexed me above all others. My posts on the subject have constituted a kind of chronicle of miseries as I thrashed back and forth trying to make sense of these things; I have made and then retracted the same decisions multiple times. To wit --

On specifiers and indefinite participant marking: First thoughts on syntax (2002), The birth of the modern article system (2003), Specifier clarification, possession, and other changes (2009), The sweeping generalization in Koa (2010), Specifier flowchart (2010), Object incorporation (2011), Marking indefinite NPs (2012), Quantifiers and specifiers (2021), What's in a specifier? (2022), Some, all and none (2025)

On dependent clauses: More Thoughts, Obtained Whilst Tossing & Turning (2007), Embedded clauses: the show so far (2012), New embedded clause options (2012), The Grand Unified Theory of dependent clauses (2019), Entr'acte: The vanity of logic and typological neutrality (2021), Resolving the probability waveform of dependent clauses (2021), Dependent clauses re-re-reenvisioned (2023), Bride of the nominalized clause (2023), Finite clause types at last (2025)

I had also been noticing that, as happy as I was in the moment with my recent progress in these areas, Koa syntax had become rigid and brittle in a way I was increasingly dismayed with. It made me less excited to work on the language or speak it, and I found myself wondering with a certain amount of resentment how it was that Nahuatl -- omnipredicative very analogously to Koa -- was allowed to be so fluid and relaxed.

And then, out of nowhere, earlier this week it suddenly dawned on me that these woes are really all the same problem. Crucially, it's not the syntactic/pragmatic design problem I thought it was: in fact, it's a cluster of dead ends manufactured largely by Indo-European grammatical terminology and the assumptions imported along with it. There are actually good answers to these formerly intractable questions, answers that are internally consistent, cross-typologically defensible and loyal to the spirit of the language, answers that let Koa sway and breathe.

Answers to other questions, sometimes surprising ones, have tumbled out alongside them this past week, and I find myself in a rather profound moment in the conceptual history of this language. All of this -- the understandings, the changes that follow from them, and a new meta-language of description that won't tie the system in knots with IE artifacts -- is still unfolding. Nonetheless I want to try to document what I now know about my own language, and where I suddenly clearly see critical mistakes to have been made in the past.

...unfortunately, or fortunately depending on how you look at it, the implications of all these changes and realizations have turned out to be quite broad. I had initially thought I would be writing a single post today that would sum it all up, but my docket has ballooned to essentially the size of a small reference grammar:

1. Lexical Class in Koa
1.1 The Particle
1.2 The Predicate
1.2.1 Predicate Functions
1.2.2 Pronominals
2. Argument Marking
2.1 Specifiers
2.2 Quantifiers
2.3 Operator Nesting / Scope layering
2.4 Compound Proforms (Correlatives)
2.5 Existential Structures
2.6 Negative Existentials & Concord
2.7. Disambiguation & Emphasis
3. Modifiers
3.1 Modifier Marking
3.2 Relativization
3.3 Object Incorporation
4. Argument Juxtaposition
4.1 Pronominal Possession
4.2 Predicate Possession
4.3 Ditransitive Arguments
4.4 Other Uses of Juxtaposition
5. Adjunctives & Relators
5.1 With Arguments
5.2 With Clauses
6. The Clause
6.1 Clause Markers
6.2 Finiteness & Nominalization
7. Constructional Specification
7.1 Categorical Typing
7.2 Output Constructions

So this is a very exciting list of topics to (A) finally understand fully and (B) get to write about, and also I wish I could take a week off of work in which to do it: I'm pretty sure the right venue for this would be a cabin by an alpine lake in Slovenia. Hopefully my girls will let me sneak away a bit this coming week.

Saturday, January 3, 2026

Koa Phonology I: Vocalism

Voa usi iolo! Viloa 2026...

Despite the fact that phonetics and phonology are arguably my favorite areas of linguistics (with honorable mention to historical change), I've managed to go 18 years on this blog saying very little about these things. It's time to make amends! This will be the first in a series in which I'll try to describe Koa's phonology with the rigor it deserves.

1. Koa Phonology

Having been conceived as an international auxiliary language, it may be helpful as a conceptual device to imagine a world in which Koa is spoken as an L2 by people from many different language backgrounds. In such a world, we would inevitably see a great deal of regional variation in the actual realization of Koa's phonology; it was in fact an original design parameter that the phonology expect and allow for this variation. Though we attempt here a description of a neutral standard, then -- coincidentally the dialect of the author -- it is important to bear in mind that no single region is intended as the language's center of gravity, validity or prestige. From a prescriptive perspective, L1 constraints may drive the precise realization of these phonemes freely as long as they remain consistent and distinct from one another.

For this reason, each section below describes both the standard and the range of expected acceptable variation we might expect to encounter around the world.

1.1 Vocalism
1.1.1 Monophthongs


Koa has the five short vowels typical of many world languages, and three long vowel phonemes; vowels are rounded only if [+back]:


Short       Long
  Front Central Back   Front Central Back
High        i [ɪ]
u [ʊ]   i: [iː]
u: [uː]
Mid e [ɛ]
o [ɔ]  


Low
a [ɑ]
 
a: [ɑː]

Long vowels are held for approximately 1.5 times the duration of short vowels. Note that mid vowels lack a long counterpart, long */e:/ and */o:/ having merged with /ei̯/ and /ou̯/ respectively (see 1.1.2 Diphthongs). Short /a/ tends to be rather more front than long /a:/. Vowel quality and quantity do not vary with stress or position.

Note: /i/ and /u/ are very peripheral and notably closer than the usual realization of the corresponding phonemes in e.g. English or German.

1.1.1.1 Variation

Provided that all phonemes remain distinct from each other, we may see variation in the following areas:

* Long vowels held for the duration of two full short vowels e.g. /a:/ → [ɑːː], or pronounced as two distinct short vowels in sequence e.g. [ɑɑ]
* Tensing of high vowels to [i u] or further laxing to [ɪ̞ ʊ̞]
* Closer mid vowels [e o]
* Frontness of /a/ anywhere from [a] to [ɑ], some rounding towards [ɒ], or some raising towards [ɐ] or [æ]
* Unrounding of back vowels, particularly /u/ as [ɯ] or [ɨ]
* Fronting of /u/ to [ʏ] or [ʉ]
* Full centralization of at most one phoneme, e.g. /a/ as [ə] or [ʌ]

Whatever the exact realization of each phoneme, it is strongly preferred that it remain consistent in quality in all positions to ease parsing by speakers of other dialects.

1.1.1.2 Devoicing

There is a tendency in rapid speech for unstressed non-initial high short vowels partially or completely to lose their voicing -- though not their duration -- following a voiceless consonant: /ákuci/ → [ˈɑkʊ̥ʃi] "hunter," /ti‿élate/ → [tɪ̥ˈɛlɑtɛ] "this life," /àpu‿taéma/ → [ˌɑpʊ̥tɑˈɛmɑ] "help his parents." This may occur even if the high vowel is stressed within its own word, but not its larger prosodic group: /ka‿somaéte‿mokùne/ → [mɔˌku̥nɛ] "what you're doing together." Utterance-final devoicing is possible but less common.

1.1.2 Diphthongs

Combinations of any two unlike vowels are possible within a single lengthened syllable. Such diphthongs exist in two types: falling (short) and balanced (long).

Falling: Where the origin is non-high and the target is high, the sequence is realized as a falling diphthong at approximately 1.2 times the duration of a short monophthongal vowel: ai [ɑɪ̯], ei [eɪ̯], oi [ɔɪ̯], au [ɑʊ̯], eu [ɛʊ̯], ou [oʊ̯].

Balanced: In other cases, diphthongs can be considered balanced in that neither the origin nor target is fully non-syllabic; the origin tends to have slightly greater duration than the target. These diphthongs are approximately the length of a long vowel, that is 1.5 times that of a short vowel: ae [ɑ͡ɛ], ao [ɑ͡ɔ], ea [ɛ͡ɑ], eo [ɛ͡ɔ], ia [ɪ͡ɑ], ie [ɪ͡ɛ], io [ɪ͡ɔ], iu [ɪ͡ʊ], oa [ɔ͡ɑ], oe [ɔ͡ɛ], ua [ʊ͡ɑ], ue [ʊ͡ɛ], ui [ʊ͡ɪ], uo [ʊ͡ɔ].

1.1.2.1 Variation

Alternative possible realizations include:

- Pronunciation of diphthongs as two consecutive full short vowels
- Diphthongs with an initial high vowel as rising: ia [ɪ̯ɑ], ie [ɪ̯ɛ], ua [ʊ̯ɑ], ue [ʊ̯ɛ], etc.
- Centralization of the origin of /ai/ and /au/ → [əɪ̯ əʊ̯], [ʌɪ̯ ʌʊ̯], [ɐɪ̯ ɐʊ̯], etc.

1.1.3 Orthography

Koa vowels are written identically to the phonemic symbols here used, whether single or in combination, except that long vowels are shown orthographically by doubling: aa /a:/ etc.

Thursday, November 27, 2025

The allowable word

Pai ko Kito iolo -- Happy Thanksgiving! After delightedly discovering eight potential new Koa predicates earlier this week that had been hiding around the edges of the phonologically permissible, I wanted to document this area more fully: what is, and is not, allowed to be a Koa word?

I've been turning a series of proper posts about Koa phonology over in my head for years, and maybe I'll finally get around to this now that the temperatures have fallen and the rains have set in. In the mean time I hope this discussion will help me firm up my understanding and description of some of the more complex corners of this topic.

So then: the simplest Koa lexical word (i.e. predicate) shapes are of the form (C)VCV, as in

CVCVhalu "want," lani "sky"
VCV: ika "acceptable," oto "crow"

The above are clearly bisyllabic, theoretically the defining criterion for predicates as opposed to particles. I say "theoretically" because another major group of Koa roots, though formally similar, behaves quite differently phonetically.

CVV: moe "dream," kai "sea," hiu "knife," paa "head," suu "mouth"

The reality is that, under ordinary circumstances, these VV sequences are not, in fact, pronounced as two consecutive full-length vowels in separate syllables, but as monosyllables: as lengthened vowels or diphthongs. We need a different phonological category here which I'm not entirely sure how to label: "long vowels" and "diphthongs" each sound exclusive of the other. "Polymoraic nuclei"? That's a bit much. Perhaps "complex vowels," against the "simple vowels" of the first two shapes?

Complex vowels, as it turns out, then, occupy about 1.2 to 1.5 times the duration of simple vowels, and thus CVV words are quite a bit shorter than CVCV words. Bisyllabic lani "sky," for example, is noticeably longer in duration than monosyllabic lai "return" or laa "therefore."

Where the two vowels are the same, we end up with a single long vowel: paa [pɑː]. This means that the orthography is imprecise at present; the aa in e.g. taahe "his arm," for instance, is pronounced as two full vowels in sequence, whereas in sáate "allowance" it represents a complex vowel of intermediate length. This distinction can probably be resolved by morphological context (i.e. sáa.te is composed of a predicate with a complex vowel followed by a suffix, whereas ta.ahe contains a predicate with two full syllables preceded by a prefix), but I have at times toyed with the idea of introducing a macron to make this absolutely explicit: taahe vs. sāte. As I suspect that my primary motivation here may simply be a love of macrons, I continue to restrain myself.

Back to the plot: where the origin and target differ, the resultant diphthong will be falling if the first member is non-high and the second is high: kai [kɑi̯] "sea," tei [tei̯] "onward," hoi [hɔi̯] "foot"; lau [lɑu̯] "shadow," ceu [ʃɛu̯] "damn," mou [mou̯] "disappear." These combinations have a relatively short duration, only slightly longer than a simple vowel. In all other combinations both vowels share syllabic value, with a transitional space between them; slight priority of duration is given to the first vowel.

I should clarify -- remembering that Koa is nominally intended as an IAL -- that it would not be incorrect to pronounce these complex vowels as sequences of two ordinary full vowels in separate syllables. In colloquial flowing speech, however, this would be unusual and would likely denote either a lack of fluency or some kind of marked pragmatic status (meli liiiiia "waaaaay too sweet").

Before continuing on, we should note one disallowed shape for predicates; these group instead with particles and affectives:

*VV: ae, ei, io, oe, ui, etc.

Now that all of the above has been thoroughly described, we can define the area of interest that originally inspired this post as a different type of predicate root: those of the form VVCV.

Right off the bat we will need to draw a distinction between two sets of what we have been calling complex vowels. This gets complicated very quickly, because vowel combinations with the particular phonetic characteristics discussed earlier do not necessarily pattern together!

On the one hand we have what we might designate as short complex vowels. These contain the true diphthongs ai, au, eu, oi -- diphthongs with one clear syllabic and one clear non-syllabic member. ei and ou, though phonetically members of this same set, pattern instead with long vowels.*

The set of short complex vowels also contains the combinations ia, ie, io, iu, because in initial or postvocalic position -- as distinct from the postconsonantal realization described earlier -- the i becomes fully non-syllabic: [ja, je, jo, ju]. Potentially similar rising dipththongs with u, however (ua, ue, ui, uo), do not pattern this way due to blocking by the phoneme /v/: as [w] is one of its primary allowable realizations, there would be no distinction between e.g. ue and ve in speakers who pronounce /v/ this way.

Our short complex vowels, then, are ai, au; eu; ia, ie, io, iu; oi.

Long complex vowels include everything else: the long vowels aa ii uu, and all remaining dipthongs ae, ao; ea, ei, eo; oa, oe, ou; ua, ue, ui, uo.

With these definitions in hand, we can now say that VVCV roots are allowed only if the initial complex vowel is short:

VVCVaimo "star," auli "willing, eager," euca "selfish," iapu "spit," iela "whole, unbroken," iolo "jolly," iuna "train," oisa "seed."

*VV:CV: aela, outo, iisi, eoku, etc. (all disallowed)

NB: /v/ is disallowed after au and eu in VVCV roots, as e.g. *euva [ɛu̯wɑ] would be all but indistinguishable from eva [ɛwɑ] for those speakers who pronounce /v/ as [w].

The last allowable predicate shape, and the object of my recent discovery mentioned above, is the sonority festival VVV. This is possible only under very specific conditions: if the initial two vowels constitute a complex vowel, if that complex vowel is short, if the resultant form would have two syllabic nuclei, if the second member of the complex vowel is not u, and if the resultant form would not violate phonological constraints. Thus:

VVV: aia, aie, aio, aiu; oia, oie, oio, oiu

In VVV roots au and eu are not allowable as the initial complex vowel because e.g. *aua would be identical in pronunciation to ava for many speakers; ia, ie, io and iu are not allowable because the resultant forms would have only one syllabic nucleus (*iuu, *iai, *ioe, etc.); the final vowel may not be /i/ because the sequence [ji] violates phonological constraints (*aii [ɑji] etc.); and forms like *eia or *uio are not allowable because the initial complex vowel is not short.

So...after all that, there are only eight! But beautiful words, which I'm very happy to have discovered.

It occurs to me that we're now in a position to define the minimum Koa lexical word as either (A) having two syllabic nuclei, or (B) having a complex syllabic nucleus with onset, in a form that violates no phonological constraints.

The maximal root word, on the other hand, might be described something like this (possibly incomplete):
1. A maximum of two syllables (*keleki)
2. If containing two syllables, neither may satisfy minimum word conditions alone (*keile, *akai)
3. If containing a long complex vowel, a maximum of one syllable (*eile)

I've asked myself about condition 3 quite a lot of times. Whereas conditions 1 and 2 serve to prevent certain and profligate misunderstanding, there's nothing objectively wrong with words like *aela...except that they're just too heavy to be Koa. The language doesn't have Estonian's oomph, and they feel like compounds even though they technically couldn't be.

...and that, ladies and gentlemen, is that! Time to go put some yams in the oven. It's been fascinating to discover how much complexity has emerged here: none of this was anything I would ever have expected in 1999.

* Sequences of identical mid vowels are disallowed within predicates, as the distinction between e.g. /ee/ and /ei/ was deemed likely to be too marginal to bear semantic load for many speakers. One could say, perhaps, that [ei̯] is the surface realization of both /ee/ and /ei/, and that [ou̯] is the surface realization of both /oo/ and /ou/. As such, ei and ou pattern with long complex vowels even though their phonetic realization, short in duration, matches that of other true diphthongs like au and oi.